task e2e-llm-inference-service has failed: "step-fail-if-needed" exited with code 1: Error [get-kubeconfig] Found kubeconfig secret: cluster-wfc8g-admin-kubeconfig [get-kubeconfig] Wrote kubeconfig to /credentials/cluster-wfc8g-kubeconfig [get-kubeconfig] Found admin password secret: cluster-wfc8g-admin-password [get-kubeconfig] Retrieved username [get-kubeconfig] Wrote password to /credentials/cluster-wfc8g-password [get-kubeconfig] API Server URL: https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443 [get-kubeconfig] Console URL: https://console-openshift-console.apps.a6ad3372-88a3-438c-8fd6-ebb9d65cfbf6.prod.konfluxeaas.com [clone-repo] fix-score-disconnected [clone-repo] https://github.com/mwaykole/kserve [clone-repo] Cloning into '/workspace/source'... [clone-repo] Updating files: 88% (2746/3093) Updating files: 89% (2753/3093) Updating files: 90% (2784/3093) Updating files: 91% (2815/3093) Updating files: 92% (2846/3093) Updating files: 93% (2877/3093) Updating files: 94% (2908/3093) Updating files: 95% (2939/3093) Updating files: 96% (2970/3093) Updating files: 97% (3001/3093) Updating files: 98% (3032/3093) Updating files: 99% (3063/3093) Updating files: 100% (3093/3093) Updating files: 100% (3093/3093), done. [e2e-llm-inference-service] + bash [e2e-llm-inference-service] + STATUS_FILE=/test-status/deploy-and-e2e-status [e2e-llm-inference-service] + echo failed [e2e-llm-inference-service] + COMPONENT_NAME=kserve-agent-ci [e2e-llm-inference-service] ++ jq -r --arg component_name kserve-agent-ci '.[$component_name].image' [e2e-llm-inference-service] + export KSERVE_AGENT_IMAGE=quay.io/opendatahub/kserve-agent@sha256:4e7b5aaca6f327ecd44367d576730108bfee7dfd141d17da265242541f7cb6f0 [e2e-llm-inference-service] + KSERVE_AGENT_IMAGE=quay.io/opendatahub/kserve-agent@sha256:4e7b5aaca6f327ecd44367d576730108bfee7dfd141d17da265242541f7cb6f0 [e2e-llm-inference-service] + COMPONENT_NAME=kserve-controller-ci [e2e-llm-inference-service] ++ jq -r --arg component_name kserve-controller-ci '.[$component_name].image' [e2e-llm-inference-service] + export KSERVE_CONTROLLER_IMAGE=quay.io/opendatahub/kserve-controller@sha256:4c416da6322c74a6fee95d39764157a74c272d597961c8986a359bfb9fb82d65 [e2e-llm-inference-service] + KSERVE_CONTROLLER_IMAGE=quay.io/opendatahub/kserve-controller@sha256:4c416da6322c74a6fee95d39764157a74c272d597961c8986a359bfb9fb82d65 [e2e-llm-inference-service] + COMPONENT_NAME=kserve-router-ci [e2e-llm-inference-service] ++ jq -r --arg component_name kserve-router-ci '.[$component_name].image' [e2e-llm-inference-service] + export KSERVE_ROUTER_IMAGE=quay.io/opendatahub/kserve-router@sha256:03660399b01ea350e32a900800c331d43d890f82d1c053ba75b58e21480c1066 [e2e-llm-inference-service] + KSERVE_ROUTER_IMAGE=quay.io/opendatahub/kserve-router@sha256:03660399b01ea350e32a900800c331d43d890f82d1c053ba75b58e21480c1066 [e2e-llm-inference-service] + COMPONENT_NAME=kserve-storage-initializer-ci [e2e-llm-inference-service] ++ jq -r --arg component_name kserve-storage-initializer-ci '.[$component_name].image' [e2e-llm-inference-service] + export STORAGE_INITIALIZER_IMAGE=quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] + STORAGE_INITIALIZER_IMAGE=quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] + COMPONENT_NAME=odh-kserve-llmisvc-controller-ci [e2e-llm-inference-service] ++ jq -r --arg component_name odh-kserve-llmisvc-controller-ci '.[$component_name].image' [e2e-llm-inference-service] + export LLMISVC_CONTROLLER_IMAGE=quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:5cf02ce74f99335a1aae18b560c417acbef30928b1c56b69e55be9ebe60b6195 [e2e-llm-inference-service] + LLMISVC_CONTROLLER_IMAGE=quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:5cf02ce74f99335a1aae18b560c417acbef30928b1c56b69e55be9ebe60b6195 [e2e-llm-inference-service] + ./test/scripts/openshift-ci/run-e2e-tests.sh 'llminferenceservice and cluster_cpu and not autoscaling and not tracing' 2 llm-d [e2e-llm-inference-service] Installing on cluster [e2e-llm-inference-service] Using namespace: kserve for KServe components [e2e-llm-inference-service] SKLEARN_IMAGE=quay.io/opendatahub/sklearn-serving-runtime:odh-pr-1694 [e2e-llm-inference-service] OPT_125M_MODEL_URI=s3://example-models/facebook/opt-125m [e2e-llm-inference-service] ERROR_404_ISVC_IMAGE=quay.io/opendatahub/error-404-isvc:odh-pr-1694 [e2e-llm-inference-service] SUCCESS_200_ISVC_IMAGE=quay.io/opendatahub/success-200-isvc:odh-pr-1694 [e2e-llm-inference-service] [INFO] Installing Kustomize v5.8.1 for linux/amd64... [e2e-llm-inference-service] [SUCCESS] Successfully installed Kustomize v5.8.1 to /workspace/source/bin/kustomize [e2e-llm-inference-service] v5.8.1 [e2e-llm-inference-service] make: Entering directory '/workspace/source' [e2e-llm-inference-service] [INFO] Installing yq v4.52.1 for linux/amd64... [e2e-llm-inference-service] [SUCCESS] Successfully installed yq v4.52.1 to /workspace/source/bin/yq [e2e-llm-inference-service] yq (https://github.com/mikefarah/yq/) version v4.52.1 [e2e-llm-inference-service] make: Leaving directory '/workspace/source' [e2e-llm-inference-service] Installing KServe Python SDK ... [e2e-llm-inference-service] [INFO] Installing uv 0.7.8 for linux/amd64... [e2e-llm-inference-service] [SUCCESS] Successfully installed uv 0.7.8 to /workspace/source/bin/uv [e2e-llm-inference-service] warning: Failed to read project metadata (No `pyproject.toml` found in current directory or any parent directory). Running `uv self version` for compatibility. This fallback will be removed in the future; pass `--preview` to force an error. [e2e-llm-inference-service] uv 0.7.8 [e2e-llm-inference-service] Creating virtual environment... [e2e-llm-inference-service] warning: virtualenv's `--clear` has no effect (uv always clears the virtual environment) [e2e-llm-inference-service] Using CPython 3.9.25 interpreter at: /usr/bin/python3 [e2e-llm-inference-service] Creating virtual environment at: .venv [e2e-llm-inference-service] /workspace/source [e2e-llm-inference-service] Using CPython 3.11.13 interpreter at: /usr/bin/python3.11 [e2e-llm-inference-service] Creating virtual environment at: .venv [e2e-llm-inference-service] Resolved 266 packages in 1ms [e2e-llm-inference-service] Building kserve @ file:///workspace/source/python/kserve [e2e-llm-inference-service] Downloading pandas (12.5MiB) [e2e-llm-inference-service] Downloading pyarrow (40.1MiB) [e2e-llm-inference-service] Downloading kubernetes (1.9MiB) [e2e-llm-inference-service] Downloading uvloop (3.8MiB) [e2e-llm-inference-service] Downloading cryptography (4.3MiB) [e2e-llm-inference-service] Downloading pydantic-core (2.0MiB) [e2e-llm-inference-service] Downloading aiohttp (1.7MiB) [e2e-llm-inference-service] Downloading mypy (17.2MiB) [e2e-llm-inference-service] Downloading botocore (12.9MiB) [e2e-llm-inference-service] Downloading numpy (15.7MiB) [e2e-llm-inference-service] Downloading setuptools (1.2MiB) [e2e-llm-inference-service] Downloading black (1.6MiB) [e2e-llm-inference-service] Downloading grpcio-tools (2.5MiB) [e2e-llm-inference-service] Downloading portforward (3.9MiB) [e2e-llm-inference-service] Downloading grpcio (6.4MiB) [e2e-llm-inference-service] Building timeout-sampler==1.0.3 [e2e-llm-inference-service] Building python-simple-logger==2.0.19 [e2e-llm-inference-service] Downloading aiohttp [e2e-llm-inference-service] Downloading black [e2e-llm-inference-service] Downloading pydantic-core [e2e-llm-inference-service] Downloading grpcio-tools [e2e-llm-inference-service] Downloading setuptools [e2e-llm-inference-service] Downloading portforward [e2e-llm-inference-service] Built python-simple-logger==2.0.19 [e2e-llm-inference-service] Downloading uvloop [e2e-llm-inference-service] Downloading cryptography [e2e-llm-inference-service] Downloading grpcio [e2e-llm-inference-service] Downloading kubernetes [e2e-llm-inference-service] Built timeout-sampler==1.0.3 [e2e-llm-inference-service] Downloading numpy [e2e-llm-inference-service] Built kserve @ file:///workspace/source/python/kserve [e2e-llm-inference-service] Downloading pandas [e2e-llm-inference-service] Downloading botocore [e2e-llm-inference-service] Downloading pyarrow [e2e-llm-inference-service] Downloading mypy [e2e-llm-inference-service] Prepared 101 packages in 1.86s [e2e-llm-inference-service] warning: Failed to hardlink files; falling back to full copy. This may lead to degraded performance. [e2e-llm-inference-service] If the cache and target directories are on different filesystems, hardlinking may not be supported. [e2e-llm-inference-service] If this is intentional, set `export UV_LINK_MODE=copy` or use `--link-mode=copy` to suppress this warning. [e2e-llm-inference-service] Installed 101 packages in 513ms [e2e-llm-inference-service] + aiohappyeyeballs==2.6.1 [e2e-llm-inference-service] + aiohttp==3.13.3 [e2e-llm-inference-service] + aiosignal==1.4.0 [e2e-llm-inference-service] + annotated-doc==0.0.4 [e2e-llm-inference-service] + annotated-types==0.7.0 [e2e-llm-inference-service] + anyio==4.9.0 [e2e-llm-inference-service] + attrs==25.3.0 [e2e-llm-inference-service] + avro==1.12.0 [e2e-llm-inference-service] + black==24.3.0 [e2e-llm-inference-service] + boto3==1.37.35 [e2e-llm-inference-service] + botocore==1.37.35 [e2e-llm-inference-service] + cachetools==5.5.2 [e2e-llm-inference-service] + certifi==2025.1.31 [e2e-llm-inference-service] + cffi==2.0.0 [e2e-llm-inference-service] + charset-normalizer==3.4.1 [e2e-llm-inference-service] + click==8.1.8 [e2e-llm-inference-service] + cloudevents==1.11.0 [e2e-llm-inference-service] + colorama==0.4.6 [e2e-llm-inference-service] + colorlog==6.10.1 [e2e-llm-inference-service] + coverage==7.8.0 [e2e-llm-inference-service] + cryptography==46.0.5 [e2e-llm-inference-service] + deprecation==2.1.0 [e2e-llm-inference-service] + durationpy==0.9 [e2e-llm-inference-service] + execnet==2.1.1 [e2e-llm-inference-service] + fastapi==0.136.3 [e2e-llm-inference-service] + frozenlist==1.5.0 [e2e-llm-inference-service] + google-auth==2.39.0 [e2e-llm-inference-service] + grpc-interceptor==0.15.4 [e2e-llm-inference-service] + grpcio==1.78.1 [e2e-llm-inference-service] + grpcio-testing==1.78.1 [e2e-llm-inference-service] + grpcio-tools==1.78.1 [e2e-llm-inference-service] + h11==0.16.0 [e2e-llm-inference-service] + httpcore==1.0.9 [e2e-llm-inference-service] + httptools==0.6.4 [e2e-llm-inference-service] + httpx==0.27.2 [e2e-llm-inference-service] + httpx-retries==0.4.5 [e2e-llm-inference-service] + idna==3.10 [e2e-llm-inference-service] + iniconfig==2.1.0 [e2e-llm-inference-service] + jinja2==3.1.6 [e2e-llm-inference-service] + jmespath==1.0.1 [e2e-llm-inference-service] + kserve==0.19.0 (from file:///workspace/source/python/kserve) [e2e-llm-inference-service] + kubernetes==32.0.1 [e2e-llm-inference-service] + markupsafe==3.0.2 [e2e-llm-inference-service] + multidict==6.4.3 [e2e-llm-inference-service] + mypy==0.991 [e2e-llm-inference-service] + mypy-extensions==1.0.0 [e2e-llm-inference-service] + numpy==2.2.4 [e2e-llm-inference-service] + oauthlib==3.2.2 [e2e-llm-inference-service] + orjson==3.10.16 [e2e-llm-inference-service] + packaging==24.2 [e2e-llm-inference-service] + pandas==2.2.3 [e2e-llm-inference-service] + pathspec==0.12.1 [e2e-llm-inference-service] + platformdirs==4.3.7 [e2e-llm-inference-service] + pluggy==1.5.0 [e2e-llm-inference-service] + portforward==0.7.1 [e2e-llm-inference-service] + prometheus-client==0.21.1 [e2e-llm-inference-service] + propcache==0.3.1 [e2e-llm-inference-service] + protobuf==6.33.5 [e2e-llm-inference-service] + psutil==5.9.8 [e2e-llm-inference-service] + pyarrow==19.0.1 [e2e-llm-inference-service] + pyasn1==0.6.3 [e2e-llm-inference-service] + pyasn1-modules==0.4.2 [e2e-llm-inference-service] + pycparser==2.22 [e2e-llm-inference-service] + pydantic==2.12.4 [e2e-llm-inference-service] + pydantic-core==2.41.5 [e2e-llm-inference-service] + pyjwt==2.12.1 [e2e-llm-inference-service] + pytest==7.4.4 [e2e-llm-inference-service] + pytest-asyncio==0.23.8 [e2e-llm-inference-service] + pytest-cov==5.0.0 [e2e-llm-inference-service] + pytest-httpx==0.30.0 [e2e-llm-inference-service] + pytest-json-report==1.5.0 [e2e-llm-inference-service] + pytest-metadata==3.1.1 [e2e-llm-inference-service] + pytest-xdist==3.6.1 [e2e-llm-inference-service] + python-dateutil==2.9.0.post0 [e2e-llm-inference-service] + python-dotenv==1.1.0 [e2e-llm-inference-service] + python-multipart==0.0.22 [e2e-llm-inference-service] + python-simple-logger==2.0.19 [e2e-llm-inference-service] + pytz==2025.2 [e2e-llm-inference-service] + pyyaml==6.0.2 [e2e-llm-inference-service] + requests==2.32.3 [e2e-llm-inference-service] + requests-oauthlib==2.0.0 [e2e-llm-inference-service] + rsa==4.9.1 [e2e-llm-inference-service] + s3transfer==0.11.4 [e2e-llm-inference-service] + setuptools==78.1.0 [e2e-llm-inference-service] + six==1.17.0 [e2e-llm-inference-service] + sniffio==1.3.1 [e2e-llm-inference-service] + starlette==1.2.1 [e2e-llm-inference-service] + tabulate==0.9.0 [e2e-llm-inference-service] + timeout-sampler==1.0.3 [e2e-llm-inference-service] + timing-asgi==0.3.1 [e2e-llm-inference-service] + tomlkit==0.13.2 [e2e-llm-inference-service] + typing-extensions==4.15.0 [e2e-llm-inference-service] + typing-inspection==0.4.2 [e2e-llm-inference-service] + tzdata==2025.2 [e2e-llm-inference-service] + urllib3==2.6.2 [e2e-llm-inference-service] + uvicorn==0.34.1 [e2e-llm-inference-service] + uvloop==0.21.0 [e2e-llm-inference-service] + watchfiles==1.0.5 [e2e-llm-inference-service] + websocket-client==1.8.0 [e2e-llm-inference-service] + websockets==15.0.1 [e2e-llm-inference-service] + yarl==1.20.0 [e2e-llm-inference-service] Audited 1 package in 50ms [e2e-llm-inference-service] /workspace/source [e2e-llm-inference-service] [INFO] Installing Kustomize v5.8.1 for linux/amd64... [e2e-llm-inference-service] [INFO] Kustomize v5.8.1 is already installed in /workspace/source/bin (>= v5.8.1) [e2e-llm-inference-service] make: Entering directory '/workspace/source' [e2e-llm-inference-service] make: Leaving directory '/workspace/source' [e2e-llm-inference-service] Now using project "kserve" on server "https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443". [e2e-llm-inference-service] [e2e-llm-inference-service] You can add applications to this project with the 'new-app' command. For example, try: [e2e-llm-inference-service] [e2e-llm-inference-service] oc new-app rails-postgresql-example [e2e-llm-inference-service] [e2e-llm-inference-service] to build a new example application in Ruby. Or use kubectl to deploy a simple Kubernetes application: [e2e-llm-inference-service] [e2e-llm-inference-service] kubectl create deployment hello-node --image=registry.k8s.io/e2e-test-images/agnhost:2.43 -- /agnhost serve-hostname [e2e-llm-inference-service] [e2e-llm-inference-service] [INFO] Installing Kustomize v5.8.1 for linux/amd64... [e2e-llm-inference-service] [INFO] Kustomize v5.8.1 is already installed in /workspace/source/bin (>= v5.8.1) [e2e-llm-inference-service] make: Entering directory '/workspace/source' [e2e-llm-inference-service] make: Leaving directory '/workspace/source' [e2e-llm-inference-service] Creating namespace openshift-keda... [e2e-llm-inference-service] namespace/openshift-keda created [e2e-llm-inference-service] Namespace openshift-keda created/ensured. [e2e-llm-inference-service] --- [e2e-llm-inference-service] Creating OperatorGroup openshift-keda... [e2e-llm-inference-service] operatorgroup.operators.coreos.com/openshift-keda created [e2e-llm-inference-service] OperatorGroup openshift-keda created/ensured. [e2e-llm-inference-service] --- [e2e-llm-inference-service] Creating Subscription for openshift-custom-metrics-autoscaler-operator... [e2e-llm-inference-service] subscription.operators.coreos.com/openshift-custom-metrics-autoscaler-operator created [e2e-llm-inference-service] Subscription openshift-custom-metrics-autoscaler-operator created/ensured. [e2e-llm-inference-service] --- [e2e-llm-inference-service] Waiting for openshift-custom-metrics-autoscaler-operator CSV to become ready... [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (0/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (5/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (10/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (15/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (20/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-custom-metrics-autoscaler-operator... (25/600) [e2e-llm-inference-service] CSV custom-metrics-autoscaler.v2.19.0-1 found, but not yet Succeeded (Phase: Installing). Waiting... (30/600) [e2e-llm-inference-service] CSV custom-metrics-autoscaler.v2.19.0-1 found, but not yet Succeeded (Phase: Installing). Waiting... (35/600) [e2e-llm-inference-service] CSV custom-metrics-autoscaler.v2.19.0-1 is ready (Phase: Succeeded). [e2e-llm-inference-service] --- [e2e-llm-inference-service] Applying KedaController custom resource... [e2e-llm-inference-service] Warning: resource kedacontrollers/keda is missing the kubectl.kubernetes.io/last-applied-configuration annotation which is required by oc apply. oc apply should only be used on resources created declaratively by either oc create --save-config or oc apply. The missing annotation will be patched automatically. [e2e-llm-inference-service] kedacontroller.keda.sh/keda configured [e2e-llm-inference-service] KedaController custom resource applied. [e2e-llm-inference-service] --- [e2e-llm-inference-service] Allowing time for KEDA components to be provisioned by the operator ... [e2e-llm-inference-service] Waiting for KEDA Operator pod (selector: "app=keda-operator") to be ready in namespace openshift-keda... [e2e-llm-inference-service] Waiting for pod -l "app=keda-operator" in namespace "openshift-keda" to be created... [e2e-llm-inference-service] Pod -l "app=keda-operator" in namespace "openshift-keda" found. [e2e-llm-inference-service] Current pods for -l "app=keda-operator" in namespace "openshift-keda": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] keda-operator-68bff59c-jv58p 1/1 Running 0 46s [e2e-llm-inference-service] Waiting up to 120s for pod(s) -l "app=keda-operator" in namespace "openshift-keda" to become ready... [e2e-llm-inference-service] pod/keda-operator-68bff59c-jv58p condition met [e2e-llm-inference-service] Pod(s) -l "app=keda-operator" in namespace "openshift-keda" are ready. [e2e-llm-inference-service] KEDA Operator pod is ready. [e2e-llm-inference-service] Waiting for KEDA Metrics API Server pod (selector: "app=keda-metrics-apiserver") to be ready in namespace openshift-keda... [e2e-llm-inference-service] Waiting for pod -l "app=keda-metrics-apiserver" in namespace "openshift-keda" to be created... [e2e-llm-inference-service] Pod -l "app=keda-metrics-apiserver" in namespace "openshift-keda" found. [e2e-llm-inference-service] Current pods for -l "app=keda-metrics-apiserver" in namespace "openshift-keda": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] keda-metrics-apiserver-559f75f947-462ps 1/1 Running 0 51s [e2e-llm-inference-service] Waiting up to 120s for pod(s) -l "app=keda-metrics-apiserver" in namespace "openshift-keda" to become ready... [e2e-llm-inference-service] pod/keda-metrics-apiserver-559f75f947-462ps condition met [e2e-llm-inference-service] Pod(s) -l "app=keda-metrics-apiserver" in namespace "openshift-keda" are ready. [e2e-llm-inference-service] KEDA Metrics API Server pod is ready. [e2e-llm-inference-service] Waiting for KEDA Webhook pod (selector: "app=keda-admission-webhooks") to be ready in namespace openshift-keda... [e2e-llm-inference-service] Waiting for pod -l "app=keda-admission-webhooks" in namespace "openshift-keda" to be created... [e2e-llm-inference-service] Pod -l "app=keda-admission-webhooks" in namespace "openshift-keda" found. [e2e-llm-inference-service] Current pods for -l "app=keda-admission-webhooks" in namespace "openshift-keda": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] keda-admission-5fbd5c4644-kcvfs 1/1 Running 0 56s [e2e-llm-inference-service] Waiting up to 120s for pod(s) -l "app=keda-admission-webhooks" in namespace "openshift-keda" to become ready... [e2e-llm-inference-service] pod/keda-admission-5fbd5c4644-kcvfs condition met [e2e-llm-inference-service] Pod(s) -l "app=keda-admission-webhooks" in namespace "openshift-keda" are ready. [e2e-llm-inference-service] KEDA Webhook pod is ready. [e2e-llm-inference-service] --- [e2e-llm-inference-service] ✅ KEDA deployment script finished successfully. [e2e-llm-inference-service] KSERVE_CONTROLLER_IMAGE=quay.io/opendatahub/kserve-controller@sha256:4c416da6322c74a6fee95d39764157a74c272d597961c8986a359bfb9fb82d65 [e2e-llm-inference-service] LLMISVC_CONTROLLER_IMAGE=quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:5cf02ce74f99335a1aae18b560c417acbef30928b1c56b69e55be9ebe60b6195 [e2e-llm-inference-service] KSERVE_AGENT_IMAGE=quay.io/opendatahub/kserve-agent@sha256:4e7b5aaca6f327ecd44367d576730108bfee7dfd141d17da265242541f7cb6f0 [e2e-llm-inference-service] KSERVE_ROUTER_IMAGE=quay.io/opendatahub/kserve-router@sha256:03660399b01ea350e32a900800c331d43d890f82d1c053ba75b58e21480c1066 [e2e-llm-inference-service] STORAGE_INITIALIZER_IMAGE=quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] Installing KServe via kustomize... [e2e-llm-inference-service] # Warning: 'commonLabels' is deprecated. Please use 'labels' instead. Run 'kustomize edit fix' to update your Kustomization automatically. [e2e-llm-inference-service] # Warning: 'commonLabels' is deprecated. Please use 'labels' instead. Run 'kustomize edit fix' to update your Kustomization automatically. [e2e-llm-inference-service] # Warning: 'commonLabels' is deprecated. Please use 'labels' instead. Run 'kustomize edit fix' to update your Kustomization automatically. [e2e-llm-inference-service] # Warning: 'commonLabels' is deprecated. Please use 'labels' instead. Run 'kustomize edit fix' to update your Kustomization automatically. [e2e-llm-inference-service] # Warning: 'commonLabels' is deprecated. Please use 'labels' instead. Run 'kustomize edit fix' to update your Kustomization automatically. [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/clusterstoragecontainers.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/datascienceclusters.datasciencecluster.opendatahub.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/dscinitializations.dscinitialization.opendatahub.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencegraphs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencemodelrewrites.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceobjectives.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepoolimports.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepools.inference.networking.k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepools.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceserviceconfigs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/servingruntimes.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/trainedmodels.serving.kserve.io serverside-applied [e2e-llm-inference-service] Waiting for CRDs to be established... [e2e-llm-inference-service] Waiting for CRD inferenceservices.serving.kserve.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD inferenceservices.serving.kserve.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceservices.serving.kserve.io condition met [e2e-llm-inference-service] Waiting for CRD llminferenceserviceconfigs.serving.kserve.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD llminferenceserviceconfigs.serving.kserve.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceserviceconfigs.serving.kserve.io condition met [e2e-llm-inference-service] Waiting for CRD clusterstoragecontainers.serving.kserve.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD clusterstoragecontainers.serving.kserve.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/clusterstoragecontainers.serving.kserve.io condition met [e2e-llm-inference-service] Waiting for CRD datascienceclusters.datasciencecluster.opendatahub.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD datascienceclusters.datasciencecluster.opendatahub.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/datascienceclusters.datasciencecluster.opendatahub.io condition met [e2e-llm-inference-service] Applying all resources... [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/clusterstoragecontainers.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/datascienceclusters.datasciencecluster.opendatahub.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/dscinitializations.dscinitialization.opendatahub.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencegraphs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencemodelrewrites.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceobjectives.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepoolimports.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepools.inference.networking.k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferencepools.inference.networking.x-k8s.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/inferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceserviceconfigs.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/llminferenceservices.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/servingruntimes.serving.kserve.io serverside-applied [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/trainedmodels.serving.kserve.io serverside-applied [e2e-llm-inference-service] serviceaccount/kserve-controller-manager serverside-applied [e2e-llm-inference-service] serviceaccount/llmisvc-controller-manager serverside-applied [e2e-llm-inference-service] role.rbac.authorization.k8s.io/kserve-leader-election-role serverside-applied [e2e-llm-inference-service] role.rbac.authorization.k8s.io/kserve-llmisvcconfig-read-access serverside-applied [e2e-llm-inference-service] role.rbac.authorization.k8s.io/llmisvc-leader-election-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-admin serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-edit serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-inferenceservice-distro-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-llmisvc-distro-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-llmisvc-manager-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-manager-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-metrics-reader-cluster-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-proxy-role serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-view serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/openshift-ai-inferenceservice-image-volume-scc serverside-applied [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/openshift-ai-llminferenceservice-scc serverside-applied [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/kserve-leader-election-rolebinding serverside-applied [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/kserve-llmisvcconfig-read-access serverside-applied [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/llmisvc-leader-election-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-inferenceservice-distro-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-llmisvc-distro-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-manager-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/kserve-proxy-rolebinding serverside-applied [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/llmisvc-manager-rolebinding serverside-applied [e2e-llm-inference-service] configmap/inferenceservice-config serverside-applied [e2e-llm-inference-service] configmap/kserve-parameters serverside-applied [e2e-llm-inference-service] secret/kserve-webhook-server-secret serverside-applied [e2e-llm-inference-service] secret/mlpipeline-s3-artifact serverside-applied [e2e-llm-inference-service] service/kserve-controller-manager-metrics-service serverside-applied [e2e-llm-inference-service] service/kserve-controller-manager-service serverside-applied [e2e-llm-inference-service] service/kserve-webhook-server-service serverside-applied [e2e-llm-inference-service] service/llmisvc-controller-manager-service serverside-applied [e2e-llm-inference-service] service/llmisvc-webhook-server-service serverside-applied [e2e-llm-inference-service] service/s3-service serverside-applied [e2e-llm-inference-service] deployment.apps/kserve-controller-manager serverside-applied [e2e-llm-inference-service] deployment.apps/llmisvc-controller-manager serverside-applied [e2e-llm-inference-service] deployment.apps/seaweedfs serverside-applied [e2e-llm-inference-service] networkpolicy.networking.k8s.io/kserve-controller-manager serverside-applied [e2e-llm-inference-service] securitycontextconstraints.security.openshift.io/openshift-ai-inferenceservice-image-volume-scc serverside-applied [e2e-llm-inference-service] securitycontextconstraints.security.openshift.io/openshift-ai-llminferenceservice-scc serverside-applied [e2e-llm-inference-service] clusterstoragecontainer.serving.kserve.io/default serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-decode-template serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-decode-worker-data-parallel serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-template serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-template-nvidia-cuda serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-template-nvidia-cuda-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-template-nvidia-cuda-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-worker-data-parallel serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-router-route serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-scheduler serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-amd-rocm serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-amd-rocm-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-amd-rocm-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-ppc64le serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-ppc64le-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-ppc64le-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-s390x serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-s390x-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-s390x-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-x86 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-x86-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-ibm-spyre-x86-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-intel-gaudi serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-intel-gaudi-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-intel-gaudi-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-nvidia-cuda serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-nvidia-cuda-fast-1 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template-nvidia-cuda-fast-2 serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-tracing serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-worker-data-parallel serverside-applied [e2e-llm-inference-service] mutatingwebhookconfiguration.admissionregistration.k8s.io/inferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] mutatingwebhookconfiguration.admissionregistration.k8s.io/llminferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/inferencegraph.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/inferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/llminferenceservice.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/llminferenceserviceconfig.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/servingruntime.serving.kserve.io serverside-applied [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/trainedmodel.serving.kserve.io serverside-applied [e2e-llm-inference-service] Waiting for llmisvc-controller-manager to be ready... [e2e-llm-inference-service] Waiting for pod -l "control-plane=llmisvc-controller-manager" in namespace "kserve" to be created... [e2e-llm-inference-service] Pod -l "control-plane=llmisvc-controller-manager" in namespace "kserve" found. [e2e-llm-inference-service] Current pods for -l "control-plane=llmisvc-controller-manager" in namespace "kserve": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] llmisvc-controller-manager-6b4f4568c4-6t9ls 0/1 Running 0 6s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "control-plane=llmisvc-controller-manager" in namespace "kserve" to become ready... [e2e-llm-inference-service] pod/llmisvc-controller-manager-6b4f4568c4-6t9ls condition met [e2e-llm-inference-service] Pod(s) -l "control-plane=llmisvc-controller-manager" in namespace "kserve" are ready. [e2e-llm-inference-service] Re-applying LLMInferenceServiceConfig resources with webhook validation... [e2e-llm-inference-service] Warning: modifying well-known config kserve/kserve-config-llm-decode-template is not recommended. Consider creating a custom config instead [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-decode-template serverside-applied [e2e-llm-inference-service] Warning: modifying well-known config kserve/kserve-config-llm-decode-worker-data-parallel is not recommended. Consider creating a custom config instead [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-decode-worker-data-parallel serverside-applied [e2e-llm-inference-service] Warning: modifying well-known config kserve/kserve-config-llm-prefill-template is not recommended. Consider creating a custom config instead [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-template serverside-applied [e2e-llm-inference-service] Warning: modifying well-known config kserve/kserve-config-llm-prefill-worker-data-parallel is not recommended. Consider creating a custom config instead [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-prefill-worker-data-parallel serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-router-route serverside-applied [e2e-llm-inference-service] Warning: modifying well-known config kserve/kserve-config-llm-scheduler is not recommended. Consider creating a custom config instead [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-scheduler serverside-applied [e2e-llm-inference-service] Warning: modifying well-known config kserve/kserve-config-llm-template is not recommended. Consider creating a custom config instead [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-template serverside-applied [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-tracing serverside-applied [e2e-llm-inference-service] Warning: modifying well-known config kserve/kserve-config-llm-worker-data-parallel is not recommended. Consider creating a custom config instead [e2e-llm-inference-service] llminferenceserviceconfig.serving.kserve.io/kserve-config-llm-worker-data-parallel serverside-applied [e2e-llm-inference-service] Applying DSC/DSCI resources... [e2e-llm-inference-service] dscinitialization.dscinitialization.opendatahub.io/test-dsci created [e2e-llm-inference-service] datasciencecluster.datasciencecluster.opendatahub.io/test-dsc created [e2e-llm-inference-service] KServe manual installation complete [e2e-llm-inference-service] 🔧 Configuration: [e2e-llm-inference-service] KServe deployment: ❌ disabled [e2e-llm-inference-service] Kuadrant deployment: ✅ enabled [e2e-llm-inference-service] [e2e-llm-inference-service] Checking OpenShift server version...(4.21.22) [e2e-llm-inference-service] 🎯 Server version (4.21.22) is 4.19.9 or higher - continue with the script [e2e-llm-inference-service] ⏳ Installing cert-manager [e2e-llm-inference-service] namespace/cert-manager-operator created [e2e-llm-inference-service] operatorgroup.operators.coreos.com/openshift-cert-manager-operator created [e2e-llm-inference-service] subscription.operators.coreos.com/openshift-cert-manager-operator created [e2e-llm-inference-service] Waiting for openshift-cert-manager-operator CSV to become ready... [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-cert-manager-operator... (0/300) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-cert-manager-operator... (5/300) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription openshift-cert-manager-operator... (10/300) [e2e-llm-inference-service] CSV cert-manager-operator.v1.19.0 found, but not yet Succeeded (Phase: Installing). Waiting... (15/300) [e2e-llm-inference-service] CSV cert-manager-operator.v1.19.0 is ready (Phase: Succeeded). [e2e-llm-inference-service] Waiting for CRD certificates.cert-manager.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD certificates.cert-manager.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/certificates.cert-manager.io condition met [e2e-llm-inference-service] ✅ Cert-manager installed [e2e-llm-inference-service] ⏳ Installing openshift-lws-operator [e2e-llm-inference-service] namespace/openshift-lws-operator created [e2e-llm-inference-service] operatorgroup.operators.coreos.com/leader-worker-set created [e2e-llm-inference-service] subscription.operators.coreos.com/leader-worker-set created [e2e-llm-inference-service] Waiting for leader-worker-set CSV to become ready... [e2e-llm-inference-service] Waiting for CSV to be installed for subscription leader-worker-set... (0/300) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription leader-worker-set... (5/300) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription leader-worker-set... (10/300) [e2e-llm-inference-service] CSV leader-worker-set.v1.0.0 found, but not yet Succeeded (Phase: Installing). Waiting... (15/300) [e2e-llm-inference-service] CSV leader-worker-set.v1.0.0 is ready (Phase: Succeeded). [e2e-llm-inference-service] Waiting for CRD leaderworkersetoperators.operator.openshift.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD leaderworkersetoperators.operator.openshift.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/leaderworkersetoperators.operator.openshift.io condition met [e2e-llm-inference-service] leaderworkersetoperator.operator.openshift.io/cluster created [e2e-llm-inference-service] ⏳ waiting for openshift-lws-operator to be ready.… [e2e-llm-inference-service] Waiting for pod -l "name=openshift-lws-operator" in namespace "openshift-lws-operator" to be created... [e2e-llm-inference-service] Pod -l "name=openshift-lws-operator" in namespace "openshift-lws-operator" found. [e2e-llm-inference-service] Current pods for -l "name=openshift-lws-operator" in namespace "openshift-lws-operator": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] openshift-lws-operator-fd8ccff4c-6fbgr 1/1 Running 0 24s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "name=openshift-lws-operator" in namespace "openshift-lws-operator" to become ready... [e2e-llm-inference-service] pod/openshift-lws-operator-fd8ccff4c-6fbgr condition met [e2e-llm-inference-service] Pod(s) -l "name=openshift-lws-operator" in namespace "openshift-lws-operator" are ready. [e2e-llm-inference-service] ✅ openshift-lws-operator installed [e2e-llm-inference-service] gatewayclass.gateway.networking.k8s.io/openshift-default created [e2e-llm-inference-service] Waiting for pod -l "app=istiod" in namespace "openshift-ingress" to be created... [e2e-llm-inference-service] Pod -l "app=istiod" in namespace "openshift-ingress" found. [e2e-llm-inference-service] Current pods for -l "app=istiod" in namespace "openshift-ingress": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] istiod-openshift-gateway-94bb8fbfd-dvdtt 0/1 Running 0 6s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "app=istiod" in namespace "openshift-ingress" to become ready... [e2e-llm-inference-service] pod/istiod-openshift-gateway-94bb8fbfd-dvdtt condition met [e2e-llm-inference-service] Pod(s) -l "app=istiod" in namespace "openshift-ingress" are ready. [e2e-llm-inference-service] ⏳ Creating a Gateway [e2e-llm-inference-service] Error from server (AlreadyExists): namespaces "openshift-ingress" already exists [e2e-llm-inference-service] gateway.gateway.networking.k8s.io/openshift-ai-inference created [e2e-llm-inference-service] Waiting for pod -l "serving.kserve.io/gateway=kserve-ingress-gateway" in namespace "openshift-ingress" to be created... [e2e-llm-inference-service] Pod -l "serving.kserve.io/gateway=kserve-ingress-gateway" in namespace "openshift-ingress" found. [e2e-llm-inference-service] Current pods for -l "serving.kserve.io/gateway=kserve-ingress-gateway" in namespace "openshift-ingress": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] openshift-ai-inference-openshift-default-6c48cf8bcc-q89vh 1/1 Running 0 5s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "serving.kserve.io/gateway=kserve-ingress-gateway" in namespace "openshift-ingress" to become ready... [e2e-llm-inference-service] pod/openshift-ai-inference-openshift-default-6c48cf8bcc-q89vh condition met [e2e-llm-inference-service] Pod(s) -l "serving.kserve.io/gateway=kserve-ingress-gateway" in namespace "openshift-ingress" are ready. [e2e-llm-inference-service] ⏳ Installing RHCL(Kuadrant) operator [e2e-llm-inference-service] namespace/kuadrant-system created [e2e-llm-inference-service] subscription.operators.coreos.com/rhcl-operator created [e2e-llm-inference-service] operatorgroup.operators.coreos.com/kuadrant created [e2e-llm-inference-service] Waiting for rhcl-operator CSV to become ready... [e2e-llm-inference-service] Waiting for CSV to be installed for subscription rhcl-operator... (0/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription rhcl-operator... (5/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription rhcl-operator... (10/600) [e2e-llm-inference-service] Waiting for CSV to be installed for subscription rhcl-operator... (15/600) [e2e-llm-inference-service] CSV rhcl-operator.v1.4.1 found, but not yet Succeeded (Phase: Installing). Waiting... (20/600) [e2e-llm-inference-service] CSV rhcl-operator.v1.4.1 found, but not yet Succeeded (Phase: Installing). Waiting... (25/600) [e2e-llm-inference-service] CSV rhcl-operator.v1.4.1 found, but not yet Succeeded (Phase: Installing). Waiting... (30/600) [e2e-llm-inference-service] CSV rhcl-operator.v1.4.1 is ready (Phase: Succeeded). [e2e-llm-inference-service] Waiting for CRD kuadrants.kuadrant.io to appear (timeout: 90s)… [e2e-llm-inference-service] CRD kuadrants.kuadrant.io detected — waiting for it to become Established (timeout: 90s)… [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/kuadrants.kuadrant.io condition met [e2e-llm-inference-service] Waiting for apiserver discovery /apis/kuadrant.io/v1beta1 to list kuadrants (timeout: 120s)… [e2e-llm-inference-service] Discovery for kuadrant.io/v1beta1 includes kuadrants. [e2e-llm-inference-service] ⏳ sleeping 30s after discovery (RESTMapper can trail discovery)… [e2e-llm-inference-service] kuadrant.kuadrant.io/kuadrant created [e2e-llm-inference-service] ⏳ waiting for Kuadrant Ready (attempt 1/2, timeout 5m)… [e2e-llm-inference-service] kuadrant.kuadrant.io/kuadrant condition met [e2e-llm-inference-service] Waiting for pod -l "control-plane=authorino-operator" in namespace "kuadrant-system" to be created... [e2e-llm-inference-service] Pod -l "control-plane=authorino-operator" in namespace "kuadrant-system" found. [e2e-llm-inference-service] Current pods for -l "control-plane=authorino-operator" in namespace "kuadrant-system": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] authorino-operator-6d85f6564-qbwjz 1/1 Running 0 76s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "control-plane=authorino-operator" in namespace "kuadrant-system" to become ready... [e2e-llm-inference-service] pod/authorino-operator-6d85f6564-qbwjz condition met [e2e-llm-inference-service] Pod(s) -l "control-plane=authorino-operator" in namespace "kuadrant-system" are ready. [e2e-llm-inference-service] ⏳ waiting for authorino service to be created... [e2e-llm-inference-service] service/authorino-authorino-authorization condition met [e2e-llm-inference-service] service/authorino-authorino-authorization annotated [e2e-llm-inference-service] Warning: resource authorinos/authorino is missing the kubectl.kubernetes.io/last-applied-configuration annotation which is required by oc apply. oc apply should only be used on resources created declaratively by either oc create --save-config or oc apply. The missing annotation will be patched automatically. [e2e-llm-inference-service] authorino.operator.authorino.kuadrant.io/authorino configured [e2e-llm-inference-service] Waiting for pod -l "control-plane=authorino-operator" in namespace "kuadrant-system" to be created... [e2e-llm-inference-service] Pod -l "control-plane=authorino-operator" in namespace "kuadrant-system" found. [e2e-llm-inference-service] Current pods for -l "control-plane=authorino-operator" in namespace "kuadrant-system": [e2e-llm-inference-service] NAME READY STATUS RESTARTS AGE [e2e-llm-inference-service] authorino-operator-6d85f6564-qbwjz 1/1 Running 0 85s [e2e-llm-inference-service] Waiting up to 600s for pod(s) -l "control-plane=authorino-operator" in namespace "kuadrant-system" to become ready... [e2e-llm-inference-service] pod/authorino-operator-6d85f6564-qbwjz condition met [e2e-llm-inference-service] Pod(s) -l "control-plane=authorino-operator" in namespace "kuadrant-system" are ready. [e2e-llm-inference-service] ✅ kuadrant(authorino) installed [e2e-llm-inference-service] Patching ingress domain... [e2e-llm-inference-service] configmap/inferenceservice-config patched [e2e-llm-inference-service] pod "kserve-controller-manager-69b5d94596-qgpch" deleted [e2e-llm-inference-service] Waiting for kserve-controller-manager to be ready... [e2e-llm-inference-service] pod/kserve-controller-manager-69b5d94596-jqzvl condition met [e2e-llm-inference-service] Installing ODH Model Controller manually... [e2e-llm-inference-service] customresourcedefinition.apiextensions.k8s.io/accounts.nim.opendatahub.io created [e2e-llm-inference-service] serviceaccount/model-serving-api created [e2e-llm-inference-service] serviceaccount/odh-model-controller created [e2e-llm-inference-service] role.rbac.authorization.k8s.io/leader-election-role created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/account-editor-role created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/account-viewer-role created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/kserve-prometheus-k8s created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/metrics-reader created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/model-serving-api created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/odh-model-controller-role created [e2e-llm-inference-service] clusterrole.rbac.authorization.k8s.io/proxy-role created [e2e-llm-inference-service] rolebinding.rbac.authorization.k8s.io/leader-election-rolebinding created [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/model-serving-api created [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/odh-model-controller-rolebinding-opendatahub created [e2e-llm-inference-service] clusterrolebinding.rbac.authorization.k8s.io/proxy-rolebinding created [e2e-llm-inference-service] configmap/odh-model-controller-parameters created [e2e-llm-inference-service] service/model-serving-api created [e2e-llm-inference-service] service/odh-model-controller-metrics-service created [e2e-llm-inference-service] service/odh-model-controller-webhook-service created [e2e-llm-inference-service] deployment.apps/model-serving-api created [e2e-llm-inference-service] deployment.apps/odh-model-controller created [e2e-llm-inference-service] servicemonitor.monitoring.coreos.com/model-serving-api-metrics created [e2e-llm-inference-service] servicemonitor.monitoring.coreos.com/odh-model-controller-metrics-monitor created [e2e-llm-inference-service] template.template.openshift.io/guardrails-detector-huggingface-serving-template created [e2e-llm-inference-service] template.template.openshift.io/kserve-ovms created [e2e-llm-inference-service] template.template.openshift.io/mlserver-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-x86-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-x86-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-cpu-x86-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-cuda-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-cuda-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-cuda-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-gaudi-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-gaudi-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-gaudi-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-multinode-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-multinode-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-multinode-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-rocm-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-rocm-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-rocm-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-ppc64le-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-ppc64le-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-ppc64le-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-s390x-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-s390x-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-s390x-runtime-template-fast-2 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-x86-runtime-template created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-x86-runtime-template-fast-1 created [e2e-llm-inference-service] template.template.openshift.io/vllm-spyre-x86-runtime-template-fast-2 created [e2e-llm-inference-service] mutatingwebhookconfiguration.admissionregistration.k8s.io/mutating.odh-model-controller.opendatahub.io created [e2e-llm-inference-service] validatingwebhookconfiguration.admissionregistration.k8s.io/validating.odh-model-controller.opendatahub.io created [e2e-llm-inference-service] Waiting for deployment "odh-model-controller" rollout to finish: 0 of 1 updated replicas are available... [e2e-llm-inference-service] deployment "odh-model-controller" successfully rolled out [e2e-llm-inference-service] networkpolicy.networking.k8s.io/allow-all created [e2e-llm-inference-service] KServe setup complete (namespace: kserve) [e2e-llm-inference-service] Add testing models to SeaweedFS S3 storage ... [e2e-llm-inference-service] Waiting for SeaweedFS deployment to be ready... [e2e-llm-inference-service] deployment "seaweedfs" successfully rolled out [e2e-llm-inference-service] S3 init job not completed, re-creating... [e2e-llm-inference-service] job.batch/s3-init replaced [e2e-llm-inference-service] Waiting for S3 init job to complete... [e2e-llm-inference-service] job.batch/s3-init condition met [e2e-llm-inference-service] Prepare CI namespace and install ServingRuntimes [e2e-llm-inference-service] Setting up CI namespace: kserve-ci-e2e-test [e2e-llm-inference-service] Tearing down CI namespace: kserve-ci-e2e-test [e2e-llm-inference-service] Namespace kserve-ci-e2e-test does not exist, skipping deletion [e2e-llm-inference-service] CI namespace teardown complete [e2e-llm-inference-service] Creating namespace kserve-ci-e2e-test [e2e-llm-inference-service] namespace/kserve-ci-e2e-test created [e2e-llm-inference-service] Applying S3 artifact secret [e2e-llm-inference-service] secret/mlpipeline-s3-artifact created [e2e-llm-inference-service] Applying storage-config secret [e2e-llm-inference-service] secret/storage-config created [e2e-llm-inference-service] Applying SeaweedFS S3 credentials secret [e2e-llm-inference-service] secret/seaweedfs-s3-creds created [e2e-llm-inference-service] Linking seaweedfs-s3-creds to default service account [e2e-llm-inference-service] Creating odh-trusted-ca-bundle configmap [e2e-llm-inference-service] configmap/odh-trusted-ca-bundle created [e2e-llm-inference-service] Installing ServingRuntimes [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-autogluonserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-huggingfaceserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-huggingfaceserver-multinode created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-lgbserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-mlserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-paddleserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-pmmlserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-predictiveserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-sklearnserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-tensorflow-serving created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-torchserve created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-tritonserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-vllmserver created [e2e-llm-inference-service] servingruntime.serving.kserve.io/kserve-xgbserver created [e2e-llm-inference-service] CI namespace setup complete [e2e-llm-inference-service] Setup complete [e2e-llm-inference-service] === E2E cluster / operator summary === [e2e-llm-inference-service] Client Version: 4.20.11 [e2e-llm-inference-service] Kustomize Version: v5.6.0 [e2e-llm-inference-service] Server Version: 4.21.22 [e2e-llm-inference-service] Kubernetes Version: v1.34.8 [e2e-llm-inference-service] ClusterVersion desired: 4.21.22 [e2e-llm-inference-service] ClusterVersion history (latest): 4.21.22 (Completed) [e2e-llm-inference-service] CSVs in kuadrant-system: [e2e-llm-inference-service] authorino-operator.v1.4.1 Succeeded [e2e-llm-inference-service] cert-manager-operator.v1.19.0 Succeeded [e2e-llm-inference-service] dns-operator.v1.4.0 Succeeded [e2e-llm-inference-service] limitador-operator.v1.4.0 Succeeded [e2e-llm-inference-service] rhcl-operator.v1.4.1 Succeeded [e2e-llm-inference-service] CSVs in openshift-keda: [e2e-llm-inference-service] authorino-operator.v1.4.1 Succeeded [e2e-llm-inference-service] cert-manager-operator.v1.19.0 Succeeded [e2e-llm-inference-service] custom-metrics-autoscaler.v2.19.0-1 Succeeded [e2e-llm-inference-service] dns-operator.v1.4.0 Succeeded [e2e-llm-inference-service] limitador-operator.v1.4.0 Succeeded [e2e-llm-inference-service] rhcl-operator.v1.4.1 Succeeded [e2e-llm-inference-service] CSVs in cert-manager-operator: [e2e-llm-inference-service] authorino-operator.v1.4.1 Succeeded [e2e-llm-inference-service] cert-manager-operator.v1.19.0 Succeeded [e2e-llm-inference-service] dns-operator.v1.4.0 Succeeded [e2e-llm-inference-service] limitador-operator.v1.4.0 Succeeded [e2e-llm-inference-service] rhcl-operator.v1.4.1 Succeeded [e2e-llm-inference-service] CSVs in openshift-lws-operator: [e2e-llm-inference-service] authorino-operator.v1.4.1 Succeeded [e2e-llm-inference-service] cert-manager-operator.v1.19.0 Succeeded [e2e-llm-inference-service] dns-operator.v1.4.0 Succeeded [e2e-llm-inference-service] leader-worker-set.v1.0.0 Succeeded [e2e-llm-inference-service] limitador-operator.v1.4.0 Succeeded [e2e-llm-inference-service] rhcl-operator.v1.4.1 Succeeded [e2e-llm-inference-service] CSVs in openshift-operators (ODH / shared operators, filtered): [e2e-llm-inference-service] authorino-operator.v1.4.1 Succeeded [e2e-llm-inference-service] dns-operator.v1.4.0 Succeeded [e2e-llm-inference-service] limitador-operator.v1.4.0 Succeeded [e2e-llm-inference-service] rhcl-operator.v1.4.1 Succeeded [e2e-llm-inference-service] Kuadrant / Authorino (diagnostics): [e2e-llm-inference-service] CRD kuadrants.kuadrant.io versions: v1beta1 served=true storage=true [e2e-llm-inference-service] Subscriptions in kuadrant-system: [e2e-llm-inference-service] authorino-operator-stable-redhat-operators-openshift-marketplace stable redhat-operators authorino-operator.v1.4.1 [e2e-llm-inference-service] dns-operator-stable-redhat-operators-openshift-marketplace stable redhat-operators dns-operator.v1.4.0 [e2e-llm-inference-service] limitador-operator-stable-redhat-operators-openshift-marketplace stable redhat-operators limitador-operator.v1.4.0 [e2e-llm-inference-service] rhcl-operator stable redhat-operators rhcl-operator.v1.4.1 [e2e-llm-inference-service] Kuadrant CR conditions (kuadrant/kuadrant-system): [e2e-llm-inference-service] Ready=True (Ready) [e2e-llm-inference-service] KServe deployments in kserve: [e2e-llm-inference-service] kserve-controller-manager: ready=1 image=quay.io/opendatahub/kserve-controller@sha256:4c416da6322c74a6fee95d39764157a74c272d597961c8986a359bfb9fb82d65 [e2e-llm-inference-service] imageID: quay.io/opendatahub/kserve-controller@sha256:4c416da6322c74a6fee95d39764157a74c272d597961c8986a359bfb9fb82d65 [e2e-llm-inference-service] odh-model-controller: ready=1 image=quay.io/opendatahub/odh-model-controller:fast [e2e-llm-inference-service] imageID: quay.io/opendatahub/odh-model-controller@sha256:1f4d2febe8e72e717f8b828958e404b656525a5cec24af58abe2a92f3e9d7e19 [e2e-llm-inference-service] llmisvc-controller-manager: ready=1 image=quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:5cf02ce74f99335a1aae18b560c417acbef30928b1c56b69e55be9ebe60b6195 [e2e-llm-inference-service] imageID: quay.io/opendatahub/odh-kserve-llmisvc-controller@sha256:5cf02ce74f99335a1aae18b560c417acbef30928b1c56b69e55be9ebe60b6195 [e2e-llm-inference-service] === End E2E cluster / operator summary === [e2e-llm-inference-service] /workspace/source [e2e-llm-inference-service] CA certificate extracted [e2e-llm-inference-service] REQUESTS_CA_BUNDLE=/tmp/ca.crt [e2e-llm-inference-service] Run E2E tests: llminferenceservice and cluster_cpu and not autoscaling and not tracing [e2e-llm-inference-service] Starting E2E functional tests ... [e2e-llm-inference-service] Parallelism requested for pytest is 2 [e2e-llm-inference-service] ============================= test session starts ============================== [e2e-llm-inference-service] platform linux -- Python 3.11.13, pytest-7.4.4, pluggy-1.5.0 -- /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] cachedir: .pytest_cache [e2e-llm-inference-service] metadata: {'Python': '3.11.13', 'Platform': 'Linux-5.14.0-427.115.1.el9_4.x86_64-x86_64-with-glibc2.34', 'Packages': {'pytest': '7.4.4', 'pluggy': '1.5.0'}, 'Plugins': {'xdist': '3.6.1', 'asyncio': '0.23.8', 'anyio': '4.9.0', 'httpx': '0.30.0', 'cov': '5.0.0', 'metadata': '3.1.1', 'json-report': '1.5.0'}, 'PLATFORM': 'el9'} [e2e-llm-inference-service] rootdir: /workspace/source/test/e2e [e2e-llm-inference-service] configfile: pytest.ini [e2e-llm-inference-service] plugins: xdist-3.6.1, asyncio-0.23.8, anyio-4.9.0, httpx-0.30.0, cov-5.0.0, metadata-3.1.1, json-report-1.5.0 [e2e-llm-inference-service] asyncio: mode=Mode.STRICT [e2e-llm-inference-service] created: 2/2 workers [e2e-llm-inference-service] 2 workers [42 items] [e2e-llm-inference-service] [e2e-llm-inference-service] scheduling tests via WorkStealingScheduling [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-precise-prefix-cache-inline-config-workload-llmd-simulator-kvcache] [e2e-llm-inference-service] llmisvc/test_gateway_section_name.py::test_gateway_section_name_propagation[cluster_single_node-cluster_cpu-with-section-name] 2026-07-02 12:52:58.860 5655 kserve INFO [conftest.py:configure_logger():40] Logger configured [e2e-llm-inference-service] 2026-07-02 12:52:58.861 5652 kserve INFO [conftest.py:configure_logger():40] Logger configured [e2e-llm-inference-service] 2026-07-02 12:52:58.875 5652 kserve.trace Checking Gateway router-gateway-1 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 12:52:58.875 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():34] Checking Gateway router-gateway-1 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 12:52:58.899 5652 kserve.trace Resource not found, creating Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-02 12:52:58.899 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():62] Resource not found, creating Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-02 12:52:58.907 5652 kserve.trace ✓ Successfully created Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-02 12:52:58.907 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():70] ✓ Successfully created Gateway router-gateway-1 [e2e-llm-inference-service] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_gateway_section_name.py::test_gateway_section_name_propagation[cluster_single_node-cluster_cpu-with-section-name] [e2e-llm-inference-service] llmisvc/test_gateway_section_name.py::test_gateway_section_name_propagation[cluster_single_node-cluster_cpu-without-section-name] 2026-07-02 12:53:24.612 5652 kserve.trace Checking Gateway router-gateway-1 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 12:53:24.612 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():34] Checking Gateway router-gateway-1 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 12:53:34.677 5652 kserve.trace ✓ Successfully updated Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-02 12:53:34.677 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():57] ✓ Successfully updated Gateway router-gateway-1 [e2e-llm-inference-service] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_gateway_section_name.py::test_gateway_section_name_propagation[cluster_single_node-cluster_cpu-without-section-name] [e2e-llm-inference-service] llmisvc/test_llm_auth.py::test_llm_auth_enabled_requires_token[cluster_cpu-cluster_single_node-auth-enabled-default] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-precise-prefix-cache-inline-config-workload-llmd-simulator-kvcache] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator0] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator0] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator1] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_auth.py::test_llm_auth_enabled_requires_token[cluster_cpu-cluster_single_node-auth-enabled-default] [e2e-llm-inference-service] llmisvc/test_llm_auth.py::test_llm_auth_invalid_token_rejected[cluster_cpu-cluster_single_node-auth-invalid-token] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator1] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator2] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_auth.py::test_llm_auth_invalid_token_rejected[cluster_cpu-cluster_single_node-auth-invalid-token] [e2e-llm-inference-service] llmisvc/test_llm_auth.py::test_llm_auth_disabled_no_token_required[cluster_cpu-cluster_single_node-auth-disabled] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator2] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf0] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_auth.py::test_llm_auth_disabled_no_token_required[cluster_cpu-cluster_single_node-auth-disabled] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-with-gateway-ref-router-with-managed-route-model-fb-opt-125m-workload-llmd-simulator] 2026-07-02 13:02:21.346 5652 kserve.trace Checking Gateway router-gateway-1 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:02:21.346 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():34] Checking Gateway router-gateway-1 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:02:21.386 5652 kserve.trace ✓ Successfully updated Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-02 13:02:21.386 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():57] ✓ Successfully updated Gateway router-gateway-1 [e2e-llm-inference-service] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf0] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf1] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-with-gateway-ref-router-with-managed-route-model-fb-opt-125m-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-custom-route-timeout-scheduler-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf1] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-pvc] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-custom-route-timeout-scheduler-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-with-refs-scheduler-managed-workload-single-cpu-model-fb-opt-125m] 2026-07-02 13:18:37.924 5652 kserve.trace Checking Gateway router-gateway-1 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:18:37.924 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():34] Checking Gateway router-gateway-1 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:18:37.970 5652 kserve.trace ✓ Successfully updated Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-02 13:18:37.970 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():57] ✓ Successfully updated Gateway router-gateway-1 [e2e-llm-inference-service] 2026-07-02 13:18:37.970 5652 kserve.trace Checking HttpRoute router-route-1 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:18:37.970 5652 kserve.trace INFO [gw_api.py:create_or_update_route():121] Checking HttpRoute router-route-1 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:18:37.973 5652 kserve.trace Resource not found, creating HttpRoute router-route-1 [e2e-llm-inference-service] 2026-07-02 13:18:37.973 5652 kserve.trace INFO [gw_api.py:create_or_update_route():149] Resource not found, creating HttpRoute router-route-1 [e2e-llm-inference-service] 2026-07-02 13:18:37.984 5652 kserve.trace ✓ Successfully created HttpRoute router-route-1 [e2e-llm-inference-service] 2026-07-02 13:18:37.984 5652 kserve.trace INFO [gw_api.py:create_or_update_route():157] ✓ Successfully created HttpRoute router-route-1 [e2e-llm-inference-service] 2026-07-02 13:18:37.984 5652 kserve.trace Checking HttpRoute router-route-2 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:18:37.984 5652 kserve.trace INFO [gw_api.py:create_or_update_route():121] Checking HttpRoute router-route-2 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:18:37.987 5652 kserve.trace Resource not found, creating HttpRoute router-route-2 [e2e-llm-inference-service] 2026-07-02 13:18:37.987 5652 kserve.trace INFO [gw_api.py:create_or_update_route():149] Resource not found, creating HttpRoute router-route-2 [e2e-llm-inference-service] 2026-07-02 13:18:37.996 5652 kserve.trace ✓ Successfully created HttpRoute router-route-2 [e2e-llm-inference-service] 2026-07-02 13:18:37.996 5652 kserve.trace INFO [gw_api.py:create_or_update_route():157] ✓ Successfully created HttpRoute router-route-2 [e2e-llm-inference-service] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-with-refs-scheduler-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-pd-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-pd-cpu-model-pvc] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-pd-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-custom-route-timeout-pd-scheduler-managed-workload-pd-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-pd-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-pvc] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-custom-route-timeout-pd-scheduler-managed-workload-pd-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-with-refs-pd-scheduler-managed-workload-pd-cpu-model-fb-opt-125m] 2026-07-02 13:31:12.791 5652 kserve.trace Checking Gateway router-gateway-2 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:31:12.791 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():34] Checking Gateway router-gateway-2 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:31:12.825 5652 kserve.trace Resource not found, creating Gateway router-gateway-2 [e2e-llm-inference-service] 2026-07-02 13:31:12.825 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():62] Resource not found, creating Gateway router-gateway-2 [e2e-llm-inference-service] 2026-07-02 13:31:12.831 5652 kserve.trace ✓ Successfully created Gateway router-gateway-2 [e2e-llm-inference-service] 2026-07-02 13:31:12.831 5652 kserve.trace INFO [gw_api.py:create_or_update_gateway():70] ✓ Successfully created Gateway router-gateway-2 [e2e-llm-inference-service] 2026-07-02 13:31:12.831 5652 kserve.trace Checking HttpRoute router-route-3 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:31:12.831 5652 kserve.trace INFO [gw_api.py:create_or_update_route():121] Checking HttpRoute router-route-3 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:31:12.834 5652 kserve.trace Resource not found, creating HttpRoute router-route-3 [e2e-llm-inference-service] 2026-07-02 13:31:12.834 5652 kserve.trace INFO [gw_api.py:create_or_update_route():149] Resource not found, creating HttpRoute router-route-3 [e2e-llm-inference-service] 2026-07-02 13:31:12.844 5652 kserve.trace ✓ Successfully created HttpRoute router-route-3 [e2e-llm-inference-service] 2026-07-02 13:31:12.844 5652 kserve.trace INFO [gw_api.py:create_or_update_route():157] ✓ Successfully created HttpRoute router-route-3 [e2e-llm-inference-service] 2026-07-02 13:31:12.844 5652 kserve.trace Checking HttpRoute router-route-4 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:31:12.844 5652 kserve.trace INFO [gw_api.py:create_or_update_route():121] Checking HttpRoute router-route-4 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] 2026-07-02 13:31:12.852 5652 kserve.trace Resource not found, creating HttpRoute router-route-4 [e2e-llm-inference-service] 2026-07-02 13:31:12.852 5652 kserve.trace INFO [gw_api.py:create_or_update_route():149] Resource not found, creating HttpRoute router-route-4 [e2e-llm-inference-service] 2026-07-02 13:31:12.860 5652 kserve.trace ✓ Successfully created HttpRoute router-route-4 [e2e-llm-inference-service] 2026-07-02 13:31:12.860 5652 kserve.trace INFO [gw_api.py:create_or_update_route():157] ✓ Successfully created HttpRoute router-route-4 [e2e-llm-inference-service] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_v1alpha1_to_v1alpha2_conversion [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_v1alpha1_to_v1alpha2_conversion [e2e-llm-inference-service] llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_v1alpha2_to_v1alpha1_conversion [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_v1alpha2_to_v1alpha1_conversion [e2e-llm-inference-service] llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_criticality_preservation_via_annotations [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_criticality_preservation_via_annotations [e2e-llm-inference-service] llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_lora_criticality_preservation [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_lora_criticality_preservation [e2e-llm-inference-service] llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_round_trip_conversion_preserves_fields [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_conversion.py::TestLLMInferenceServiceConversion::test_round_trip_conversion_preserves_fields [e2e-llm-inference-service] llmisvc/test_llm_inference_service_stop.py::test_llm_stop_feature[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-with-refs-pd-scheduler-managed-workload-pd-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-no-scheduler-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw1] PASSED llmisvc/test_llm_inference_service_stop.py::test_llm_stop_feature[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_lora_adapters.py::test_llm_with_lora_adapters[cluster_cpu-single-lora-adapter-hf] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-no-scheduler-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw0] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-inline-config-workload-llmd-simulator] [e2e-llm-inference-service] [gw1] FAILED llmisvc/test_llm_lora_adapters.py::test_llm_with_lora_adapters[cluster_cpu-single-lora-adapter-hf] [e2e-llm-inference-service] llmisvc/test_llm_lora_adapters.py::test_llm_with_lora_adapters[cluster_cpu-multiple-lora-adapters] [e2e-llm-inference-service] [gw0] PASSED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-inline-config-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-model-qwen2.5-0.5b] [e2e-llm-inference-service] [gw1] FAILED llmisvc/test_llm_lora_adapters.py::test_llm_with_lora_adapters[cluster_cpu-multiple-lora-adapters] [e2e-llm-inference-service] llmisvc/test_llm_tls.py::test_llm_tls_resources[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw0] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-model-qwen2.5-0.5b] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-configmap-ref-workload-llmd-simulator] [e2e-llm-inference-service] [gw1] FAILED llmisvc/test_llm_tls.py::test_llm_tls_resources[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_prestop_hook.py::test_prestop_hook[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] [gw0] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-configmap-ref-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-replicas-workload-llmd-simulator] [e2e-llm-inference-service] [gw0] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-replicas-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-custom-template-workload-llmd-simulator] [e2e-llm-inference-service] [gw0] ERROR llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-custom-template-workload-llmd-simulator] [e2e-llm-inference-service] [gw1] FAILED llmisvc/test_prestop_hook.py::test_prestop_hook[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_rolling_upgrade.py::test_rolling_upgrade_coordination[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-model-fb-opt-125m] [e2e-llm-inference-service] [gw1] ERROR llmisvc/test_rolling_upgrade.py::test_rolling_upgrade_coordination[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-model-fb-opt-125m] [e2e-llm-inference-service] [e2e-llm-inference-service] ==================================== ERRORS ==================================== [e2e-llm-inference-service] _ ERROR at setup of test_llm_inference_service[router-managed-scheduler-with-custom-template-workload-llmd-simulator] _ [e2e-llm-inference-service] [gw0] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] llm_config = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...ustom-1626c793', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}} [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test' [e2e-llm-inference-service] [e2e-llm-inference-service] def _create_or_update_llmisvc_config(kserve_client, llm_config, namespace=None): [e2e-llm-inference-service] """Create or update an LLMInferenceServiceConfig resource.""" [e2e-llm-inference-service] version = llm_config["apiVersion"].split("/")[1] [e2e-llm-inference-service] [e2e-llm-inference-service] if namespace is None: [e2e-llm-inference-service] namespace = llm_config.get("metadata", {}).get("namespace", "default") [e2e-llm-inference-service] [e2e-llm-inference-service] name = llm_config.get("metadata", {}).get("name") [e2e-llm-inference-service] if not name: [e2e-llm-inference-service] raise ValueError("LLMInferenceServiceConfig must have a name in metadata") [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Checking LLMInferenceServiceConfig {name} in namespace {namespace}") [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > existing_config = kserve_client.api_instance.get_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICECONFIG, [e2e-llm-inference-service] name, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1589: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test', plural = 'llminferenceserviceconfigs' [e2e-llm-inference-service] name = 'router-managed-scheduler-custom-1626c793' [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] [e2e-llm-inference-service] def get_namespaced_custom_object(self, group, version, namespace, plural, name, **kwargs): # noqa: E501 [e2e-llm-inference-service] """get_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a namespace scoped custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.get_namespaced_custom_object(group, version, namespace, plural, name, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: the custom resource's group (required) [e2e-llm-inference-service] :param str version: the custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: the custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param str name: the custom object's name (required) [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: object [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] kwargs['_return_http_data_only'] = True [e2e-llm-inference-service] > return self.get_namespaced_custom_object_with_http_info(group, version, namespace, plural, name, **kwargs) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:1632: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test', plural = 'llminferenceserviceconfigs' [e2e-llm-inference-service] name = 'router-managed-scheduler-custom-1626c793' [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] local_var_params = {'_return_http_data_only': True, 'all_params': ['group', 'version', 'namespace', 'plural', 'name', 'async_req', ...], 'auth_settings': ['BearerToken'], 'body_params': None, ...} [e2e-llm-inference-service] all_params = ['group', 'version', 'namespace', 'plural', 'name', 'async_req', ...] [e2e-llm-inference-service] key = '_return_http_data_only', val = True, collection_formats = {} [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'name': 'router-managed-scheduler-custom-1626c793', 'namespace': 'kserve-ci-e2e-test', 'plural': 'llminferenceserviceconfigs', ...} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] def get_namespaced_custom_object_with_http_info(self, group, version, namespace, plural, name, **kwargs): # noqa: E501 [e2e-llm-inference-service] """get_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a namespace scoped custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.get_namespaced_custom_object_with_http_info(group, version, namespace, plural, name, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: the custom resource's group (required) [e2e-llm-inference-service] :param str version: the custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: the custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param str name: the custom object's name (required) [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: tuple(object, status_code(int), headers(HTTPHeaderDict)) [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] local_var_params = locals() [e2e-llm-inference-service] [e2e-llm-inference-service] all_params = [ [e2e-llm-inference-service] 'group', [e2e-llm-inference-service] 'version', [e2e-llm-inference-service] 'namespace', [e2e-llm-inference-service] 'plural', [e2e-llm-inference-service] 'name' [e2e-llm-inference-service] ] [e2e-llm-inference-service] all_params.extend( [e2e-llm-inference-service] [ [e2e-llm-inference-service] 'async_req', [e2e-llm-inference-service] '_return_http_data_only', [e2e-llm-inference-service] '_preload_content', [e2e-llm-inference-service] '_request_timeout' [e2e-llm-inference-service] ] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] for key, val in six.iteritems(local_var_params['kwargs']): [e2e-llm-inference-service] if key not in all_params: [e2e-llm-inference-service] raise ApiTypeError( [e2e-llm-inference-service] "Got an unexpected keyword argument '%s'" [e2e-llm-inference-service] " to method get_namespaced_custom_object" % key [e2e-llm-inference-service] ) [e2e-llm-inference-service] local_var_params[key] = val [e2e-llm-inference-service] del local_var_params['kwargs'] [e2e-llm-inference-service] # verify the required parameter 'group' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('group' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['group'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `group` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'version' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('version' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['version'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `version` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'namespace' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('namespace' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['namespace'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `namespace` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'plural' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('plural' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['plural'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `plural` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'name' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('name' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['name'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `name` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] collection_formats = {} [e2e-llm-inference-service] [e2e-llm-inference-service] path_params = {} [e2e-llm-inference-service] if 'group' in local_var_params: [e2e-llm-inference-service] path_params['group'] = local_var_params['group'] # noqa: E501 [e2e-llm-inference-service] if 'version' in local_var_params: [e2e-llm-inference-service] path_params['version'] = local_var_params['version'] # noqa: E501 [e2e-llm-inference-service] if 'namespace' in local_var_params: [e2e-llm-inference-service] path_params['namespace'] = local_var_params['namespace'] # noqa: E501 [e2e-llm-inference-service] if 'plural' in local_var_params: [e2e-llm-inference-service] path_params['plural'] = local_var_params['plural'] # noqa: E501 [e2e-llm-inference-service] if 'name' in local_var_params: [e2e-llm-inference-service] path_params['name'] = local_var_params['name'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] header_params = {} [e2e-llm-inference-service] [e2e-llm-inference-service] form_params = [] [e2e-llm-inference-service] local_var_files = {} [e2e-llm-inference-service] [e2e-llm-inference-service] body_params = None [e2e-llm-inference-service] # HTTP header `Accept` [e2e-llm-inference-service] header_params['Accept'] = self.api_client.select_header_accept( [e2e-llm-inference-service] ['application/json']) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] # Authentication setting [e2e-llm-inference-service] auth_settings = ['BearerToken'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.api_client.call_api( [e2e-llm-inference-service] '/apis/{group}/{version}/namespaces/{namespace}/{plural}/{name}', 'GET', [e2e-llm-inference-service] path_params, [e2e-llm-inference-service] query_params, [e2e-llm-inference-service] header_params, [e2e-llm-inference-service] body=body_params, [e2e-llm-inference-service] post_params=form_params, [e2e-llm-inference-service] files=local_var_files, [e2e-llm-inference-service] response_type='object', # noqa: E501 [e2e-llm-inference-service] auth_settings=auth_settings, [e2e-llm-inference-service] async_req=local_var_params.get('async_req'), [e2e-llm-inference-service] _return_http_data_only=local_var_params.get('_return_http_data_only'), # noqa: E501 [e2e-llm-inference-service] _preload_content=local_var_params.get('_preload_content', True), [e2e-llm-inference-service] _request_timeout=local_var_params.get('_request_timeout'), [e2e-llm-inference-service] collection_formats=collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:1739: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/{group}/{version}/namespaces/{namespace}/{plural}/{name}' [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'name': 'router-managed-scheduler-custom-1626c793', 'namespace': 'kserve-ci-e2e-test', 'plural': 'llminferenceserviceconfigs', ...} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], async_req = None, _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def call_api(self, resource_path, method, [e2e-llm-inference-service] path_params=None, query_params=None, header_params=None, [e2e-llm-inference-service] body=None, post_params=None, files=None, [e2e-llm-inference-service] response_type=None, auth_settings=None, async_req=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] """Makes the HTTP request (synchronous) and returns deserialized data. [e2e-llm-inference-service] [e2e-llm-inference-service] To make an async_req request, set the async_req parameter. [e2e-llm-inference-service] [e2e-llm-inference-service] :param resource_path: Path to method endpoint. [e2e-llm-inference-service] :param method: Method to call. [e2e-llm-inference-service] :param path_params: Path parameters in the url. [e2e-llm-inference-service] :param query_params: Query parameters in the url. [e2e-llm-inference-service] :param header_params: Header parameters to be [e2e-llm-inference-service] placed in the request header. [e2e-llm-inference-service] :param body: Request body. [e2e-llm-inference-service] :param post_params dict: Request post form parameters, [e2e-llm-inference-service] for `application/x-www-form-urlencoded`, `multipart/form-data`. [e2e-llm-inference-service] :param auth_settings list: Auth Settings names for the request. [e2e-llm-inference-service] :param response: Response data type. [e2e-llm-inference-service] :param files dict: key -> filename, value -> filepath, [e2e-llm-inference-service] for `multipart/form-data`. [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param collection_formats: dict of collection formats for path, query, [e2e-llm-inference-service] header, and post parameters. [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: [e2e-llm-inference-service] If async_req parameter is True, [e2e-llm-inference-service] the request will be called asynchronously. [e2e-llm-inference-service] The method will return the request thread. [e2e-llm-inference-service] If parameter async_req is False or missing, [e2e-llm-inference-service] then the method will return the response directly. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if not async_req: [e2e-llm-inference-service] > return self.__call_api(resource_path, method, [e2e-llm-inference-service] path_params, query_params, header_params, [e2e-llm-inference-service] body, post_params, files, [e2e-llm-inference-service] response_type, auth_settings, [e2e-llm-inference-service] _return_http_data_only, collection_formats, [e2e-llm-inference-service] _preload_content, _request_timeout, _host) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:348: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-scheduler-custom-1626c793' [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] path_params = [('group', 'serving.kserve.io'), ('version', 'v1alpha1'), ('namespace', 'kserve-ci-e2e-test'), ('plural', 'llminferenceserviceconfigs'), ('name', 'router-managed-scheduler-custom-1626c793')] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def __call_api( [e2e-llm-inference-service] self, resource_path, method, path_params=None, [e2e-llm-inference-service] query_params=None, header_params=None, body=None, post_params=None, [e2e-llm-inference-service] files=None, response_type=None, auth_settings=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] [e2e-llm-inference-service] config = self.configuration [e2e-llm-inference-service] [e2e-llm-inference-service] # header parameters [e2e-llm-inference-service] header_params = header_params or {} [e2e-llm-inference-service] header_params.update(self.default_headers) [e2e-llm-inference-service] if self.cookie: [e2e-llm-inference-service] header_params['Cookie'] = self.cookie [e2e-llm-inference-service] if header_params: [e2e-llm-inference-service] header_params = self.sanitize_for_serialization(header_params) [e2e-llm-inference-service] header_params = dict(self.parameters_to_tuples(header_params, [e2e-llm-inference-service] collection_formats)) [e2e-llm-inference-service] [e2e-llm-inference-service] # path parameters [e2e-llm-inference-service] if path_params: [e2e-llm-inference-service] path_params = self.sanitize_for_serialization(path_params) [e2e-llm-inference-service] path_params = self.parameters_to_tuples(path_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] for k, v in path_params: [e2e-llm-inference-service] # specified safe chars, encode everything [e2e-llm-inference-service] resource_path = resource_path.replace( [e2e-llm-inference-service] '{%s}' % k, [e2e-llm-inference-service] quote(str(v), safe=config.safe_chars_for_path_param) [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # query parameters [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] query_params = self.sanitize_for_serialization(query_params) [e2e-llm-inference-service] query_params = self.parameters_to_tuples(query_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] # post parameters [e2e-llm-inference-service] if post_params or files: [e2e-llm-inference-service] post_params = post_params if post_params else [] [e2e-llm-inference-service] post_params = self.sanitize_for_serialization(post_params) [e2e-llm-inference-service] post_params = self.parameters_to_tuples(post_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] post_params.extend(self.files_parameters(files)) [e2e-llm-inference-service] [e2e-llm-inference-service] # auth setting [e2e-llm-inference-service] self.update_params_for_auth(header_params, query_params, auth_settings) [e2e-llm-inference-service] [e2e-llm-inference-service] # body [e2e-llm-inference-service] if body: [e2e-llm-inference-service] body = self.sanitize_for_serialization(body) [e2e-llm-inference-service] [e2e-llm-inference-service] # request url [e2e-llm-inference-service] if _host is None: [e2e-llm-inference-service] url = self.configuration.host + resource_path [e2e-llm-inference-service] else: [e2e-llm-inference-service] # use server/host defined in path or operation instead [e2e-llm-inference-service] url = _host + resource_path [e2e-llm-inference-service] [e2e-llm-inference-service] # perform request and return response [e2e-llm-inference-service] > response_data = self.request( [e2e-llm-inference-service] method, url, query_params=query_params, headers=header_params, [e2e-llm-inference-service] post_params=post_params, body=body, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:180: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-scheduler-custom-1626c793' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] post_params = [], body = None, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] post_params=None, body=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Makes the HTTP request using RESTClient.""" [e2e-llm-inference-service] if method == "GET": [e2e-llm-inference-service] > return self.rest_client.GET(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:373: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-scheduler-custom-1626c793' [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] query_params = [], _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def GET(self, url, headers=None, query_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] > return self.request("GET", url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] query_params=query_params) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:244: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-scheduler-custom-1626c793' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] body=None, post_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Perform requests. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: http request method [e2e-llm-inference-service] :param url: http request url [e2e-llm-inference-service] :param query_params: query parameters in the url [e2e-llm-inference-service] :param headers: http request headers [e2e-llm-inference-service] :param body: request json body, for `application/json` [e2e-llm-inference-service] :param post_params: request post parameters, [e2e-llm-inference-service] `application/x-www-form-urlencoded` [e2e-llm-inference-service] and `multipart/form-data` [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] """ [e2e-llm-inference-service] method = method.upper() [e2e-llm-inference-service] assert method in ['GET', 'HEAD', 'DELETE', 'POST', 'PUT', [e2e-llm-inference-service] 'PATCH', 'OPTIONS'] [e2e-llm-inference-service] [e2e-llm-inference-service] if post_params and body: [e2e-llm-inference-service] raise ApiValueError( [e2e-llm-inference-service] "body parameter cannot be used with post_params parameter." [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] post_params = post_params or {} [e2e-llm-inference-service] headers = headers or {} [e2e-llm-inference-service] [e2e-llm-inference-service] timeout = None [e2e-llm-inference-service] if _request_timeout: [e2e-llm-inference-service] if isinstance(_request_timeout, (int, ) if six.PY3 else (int, long)): # noqa: E501,F821 [e2e-llm-inference-service] timeout = urllib3.Timeout(total=_request_timeout) [e2e-llm-inference-service] elif (isinstance(_request_timeout, tuple) and [e2e-llm-inference-service] len(_request_timeout) == 2): [e2e-llm-inference-service] timeout = urllib3.Timeout( [e2e-llm-inference-service] connect=_request_timeout[0], read=_request_timeout[1]) [e2e-llm-inference-service] [e2e-llm-inference-service] if 'Content-Type' not in headers: [e2e-llm-inference-service] headers['Content-Type'] = 'application/json' [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # For `POST`, `PUT`, `PATCH`, `OPTIONS`, `DELETE` [e2e-llm-inference-service] if method in ['POST', 'PUT', 'PATCH', 'OPTIONS', 'DELETE']: [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] url += '?' + urlencode(query_params) [e2e-llm-inference-service] if (re.search('json', headers['Content-Type'], re.IGNORECASE) or [e2e-llm-inference-service] headers['Content-Type'] == 'application/apply-patch+yaml'): [e2e-llm-inference-service] if headers['Content-Type'] == 'application/json-patch+json': [e2e-llm-inference-service] if not isinstance(body, list): [e2e-llm-inference-service] headers['Content-Type'] = \ [e2e-llm-inference-service] 'application/strategic-merge-patch+json' [e2e-llm-inference-service] request_body = None [e2e-llm-inference-service] if body is not None: [e2e-llm-inference-service] request_body = json.dumps(body) [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'application/x-www-form-urlencoded': # noqa: E501 [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=False, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'multipart/form-data': [e2e-llm-inference-service] # must del headers['Content-Type'], or the correct [e2e-llm-inference-service] # Content-Type which generated by urllib3 will be [e2e-llm-inference-service] # overwritten. [e2e-llm-inference-service] del headers['Content-Type'] [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=True, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] # Pass a `string` parameter directly in the body to support [e2e-llm-inference-service] # other content types than Json when `body` argument is [e2e-llm-inference-service] # provided in serialized form [e2e-llm-inference-service] elif isinstance(body, str) or isinstance(body, bytes): [e2e-llm-inference-service] request_body = body [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Cannot generate the request from given parameters [e2e-llm-inference-service] msg = """Cannot prepare a request message for provided [e2e-llm-inference-service] arguments. Please check that your arguments match [e2e-llm-inference-service] declared content type.""" [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] # For `GET`, `HEAD` [e2e-llm-inference-service] else: [e2e-llm-inference-service] r = self.pool_manager.request(method, url, [e2e-llm-inference-service] fields=query_params, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] except urllib3.exceptions.SSLError as e: [e2e-llm-inference-service] msg = "{0}\n{1}".format(type(e).__name__, str(e)) [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] [e2e-llm-inference-service] if _preload_content: [e2e-llm-inference-service] r = RESTResponse(r) [e2e-llm-inference-service] [e2e-llm-inference-service] # In the python 3, the response.data is bytes. [e2e-llm-inference-service] # we need to decode it to string. [e2e-llm-inference-service] if six.PY3: [e2e-llm-inference-service] r.data = r.data.decode('utf8') [e2e-llm-inference-service] [e2e-llm-inference-service] # log response body [e2e-llm-inference-service] logger.debug("response body: %s", r.data) [e2e-llm-inference-service] [e2e-llm-inference-service] if not 200 <= r.status <= 299: [e2e-llm-inference-service] > raise ApiException(http_resp=r) [e2e-llm-inference-service] E kubernetes.client.exceptions.ApiException: (404) [e2e-llm-inference-service] E Reason: Not Found [e2e-llm-inference-service] E HTTP response headers: HTTPHeaderDict({'Audit-Id': '8176a5ad-5f37-41f8-baaa-6803426aef4c', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': 'a8664215-74a2-48d8-b231-e657e37b3200', 'X-Kubernetes-Pf-Prioritylevel-Uid': '0e7454c5-cfcf-4e4f-b47e-e88e460e3769', 'Date': 'Thu, 02 Jul 2026 14:29:47 GMT', 'Content-Length': '338'}) [e2e-llm-inference-service] E HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"llminferenceserviceconfigs.serving.kserve.io \"router-managed-scheduler-custom-1626c793\" not found","reason":"NotFound","details":{"name":"router-managed-scheduler-custom-1626c793","group":"serving.kserve.io","kind":"llminferenceserviceconfigs"},"code":404} [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:238: ApiException [e2e-llm-inference-service] [e2e-llm-inference-service] During handling of the above exception, another exception occurred: [e2e-llm-inference-service] [e2e-llm-inference-service] request = > [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.fixture(scope="function") [e2e-llm-inference-service] def test_case(request): [e2e-llm-inference-service] tc = request.param [e2e-llm-inference-service] [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Execute before test hooks [e2e-llm-inference-service] try: [e2e-llm-inference-service] for func in tc.before_test: [e2e-llm-inference-service] func() [e2e-llm-inference-service] except Exception as before_test_error: [e2e-llm-inference-service] raise RuntimeError( [e2e-llm-inference-service] f"Failed to execute before test hook: {before_test_error}" [e2e-llm-inference-service] ) from before_test_error [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > _setup_test_case_service(kserve_client, tc, request.node.name) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1476: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] tc = TestCase(base_refs=['router-managed', 'scheduler-with-custom-template', 'workload-llmd-simulator'], prompt='KServe is ...None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service=None, model_name='facebook/opt-125m') [e2e-llm-inference-service] test_node_name = 'test_llm_inference_service[router-managed-scheduler-with-custom-template-workload-llmd-simulator]' [e2e-llm-inference-service] peer_index = None [e2e-llm-inference-service] [e2e-llm-inference-service] def _setup_test_case_service(kserve_client, tc, test_node_name, peer_index=None): [e2e-llm-inference-service] """Create LLMInferenceServiceConfigs and build the LLMInferenceService for a TestCase. [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a list of created config names for cleanup tracking. [e2e-llm-inference-service] """ [e2e-llm-inference-service] missing_refs = [ [e2e-llm-inference-service] ref for ref in tc.base_refs if ref not in LLMINFERENCESERVICE_CONFIGS [e2e-llm-inference-service] ] [e2e-llm-inference-service] if missing_refs: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Missing base_refs in LLMINFERENCESERVICE_CONFIGS: {missing_refs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] if not tc.service_name: [e2e-llm-inference-service] suffix = f"-peer-{peer_index}" if peer_index is not None else "" [e2e-llm-inference-service] tc.service_name = generate_service_name(test_node_name + suffix, tc.base_refs) [e2e-llm-inference-service] if tc.model_name == "default/model": [e2e-llm-inference-service] tc.model_name = _get_model_name_from_configs(tc.base_refs) [e2e-llm-inference-service] [e2e-llm-inference-service] created_configs = [] [e2e-llm-inference-service] unique_base_refs = [] [e2e-llm-inference-service] for base_ref in tc.base_refs: [e2e-llm-inference-service] unique_config_name = generate_k8s_safe_suffix(base_ref, [tc.service_name]) [e2e-llm-inference-service] unique_base_refs.append(unique_config_name) [e2e-llm-inference-service] [e2e-llm-inference-service] unique_config_body = { [e2e-llm-inference-service] "apiVersion": "serving.kserve.io/v1alpha1", [e2e-llm-inference-service] "kind": "LLMInferenceServiceConfig", [e2e-llm-inference-service] "metadata": { [e2e-llm-inference-service] "name": unique_config_name, [e2e-llm-inference-service] "namespace": KSERVE_TEST_NAMESPACE, [e2e-llm-inference-service] }, [e2e-llm-inference-service] "spec": LLMINFERENCESERVICE_CONFIGS[base_ref], [e2e-llm-inference-service] } [e2e-llm-inference-service] [e2e-llm-inference-service] > _create_or_update_llmisvc_config( [e2e-llm-inference-service] kserve_client, unique_config_body, KSERVE_TEST_NAMESPACE [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1436: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] llm_config = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...ustom-1626c793', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}} [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test' [e2e-llm-inference-service] [e2e-llm-inference-service] def _create_or_update_llmisvc_config(kserve_client, llm_config, namespace=None): [e2e-llm-inference-service] """Create or update an LLMInferenceServiceConfig resource.""" [e2e-llm-inference-service] version = llm_config["apiVersion"].split("/")[1] [e2e-llm-inference-service] [e2e-llm-inference-service] if namespace is None: [e2e-llm-inference-service] namespace = llm_config.get("metadata", {}).get("namespace", "default") [e2e-llm-inference-service] [e2e-llm-inference-service] name = llm_config.get("metadata", {}).get("name") [e2e-llm-inference-service] if not name: [e2e-llm-inference-service] raise ValueError("LLMInferenceServiceConfig must have a name in metadata") [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Checking LLMInferenceServiceConfig {name} in namespace {namespace}") [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] existing_config = kserve_client.api_instance.get_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICECONFIG, [e2e-llm-inference-service] name, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llm_config["metadata"] = existing_config["metadata"] [e2e-llm-inference-service] [e2e-llm-inference-service] outputs = kserve_client.api_instance.replace_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICECONFIG, [e2e-llm-inference-service] name, [e2e-llm-inference-service] llm_config, [e2e-llm-inference-service] ) [e2e-llm-inference-service] logger.info(f"✓ Successfully updated LLMInferenceServiceConfig {name}") [e2e-llm-inference-service] return outputs [e2e-llm-inference-service] [e2e-llm-inference-service] except client.rest.ApiException as e: [e2e-llm-inference-service] if e.status == 404: # Not found - create it [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"Resource not found, creating LLMInferenceServiceConfig {name}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] > outputs = kserve_client.api_instance.create_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICECONFIG, [e2e-llm-inference-service] llm_config, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1615: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test', plural = 'llminferenceserviceconfigs' [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...ustom-1626c793', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}} [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] [e2e-llm-inference-service] def create_namespaced_custom_object(self, group, version, namespace, plural, body, **kwargs): # noqa: E501 [e2e-llm-inference-service] """create_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Creates a namespace scoped Custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.create_namespaced_custom_object(group, version, namespace, plural, body, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: The custom resource's group name (required) [e2e-llm-inference-service] :param str version: The custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: The custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param object body: The JSON schema of the Resource to create. (required) [e2e-llm-inference-service] :param str pretty: If 'true', then the output is pretty printed. [e2e-llm-inference-service] :param str dry_run: When present, indicates that modifications should not be persisted. An invalid or unrecognized dryRun directive will result in an error response and no further processing of the request. Valid values are: - All: all dry run stages will be processed [e2e-llm-inference-service] :param str field_manager: fieldManager is a name associated with the actor or entity that is making these changes. The value must be less than or 128 characters long, and only contain printable characters, as defined by https://golang.org/pkg/unicode/#IsPrint. [e2e-llm-inference-service] :param str field_validation: fieldValidation instructs the server on how to handle objects in the request (POST/PUT/PATCH) containing unknown or duplicate fields. Valid values are: - Ignore: This will ignore any unknown fields that are silently dropped from the object, and will ignore all but the last duplicate field that the decoder encounters. This is the default behavior prior to v1.23. - Warn: This will send a warning via the standard warning response header for each unknown field that is dropped from the object, and for each duplicate field that is encountered. The request will still succeed if there are no other errors, and will only persist the last of any duplicate fields. This is the default in v1.23+ - Strict: This will fail the request with a BadRequest error if any unknown fields would be dropped from the object, or if any duplicate fields are present. The error returned from the server will contain all unknown and duplicate fields encountered. (optional) [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: object [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] kwargs['_return_http_data_only'] = True [e2e-llm-inference-service] > return self.create_namespaced_custom_object_with_http_info(group, version, namespace, plural, body, **kwargs) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:231: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test', plural = 'llminferenceserviceconfigs' [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...ustom-1626c793', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}} [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] local_var_params = {'_return_http_data_only': True, 'all_params': ['group', 'version', 'namespace', 'plural', 'body', 'pretty', ...], 'au...1626c793', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}}, ...} [e2e-llm-inference-service] all_params = ['group', 'version', 'namespace', 'plural', 'body', 'pretty', ...] [e2e-llm-inference-service] key = '_return_http_data_only', val = True, collection_formats = {} [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'namespace': 'kserve-ci-e2e-test', 'plural': 'llminferenceserviceconfigs', 'version': 'v1alpha1'} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] def create_namespaced_custom_object_with_http_info(self, group, version, namespace, plural, body, **kwargs): # noqa: E501 [e2e-llm-inference-service] """create_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Creates a namespace scoped Custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.create_namespaced_custom_object_with_http_info(group, version, namespace, plural, body, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: The custom resource's group name (required) [e2e-llm-inference-service] :param str version: The custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: The custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param object body: The JSON schema of the Resource to create. (required) [e2e-llm-inference-service] :param str pretty: If 'true', then the output is pretty printed. [e2e-llm-inference-service] :param str dry_run: When present, indicates that modifications should not be persisted. An invalid or unrecognized dryRun directive will result in an error response and no further processing of the request. Valid values are: - All: all dry run stages will be processed [e2e-llm-inference-service] :param str field_manager: fieldManager is a name associated with the actor or entity that is making these changes. The value must be less than or 128 characters long, and only contain printable characters, as defined by https://golang.org/pkg/unicode/#IsPrint. [e2e-llm-inference-service] :param str field_validation: fieldValidation instructs the server on how to handle objects in the request (POST/PUT/PATCH) containing unknown or duplicate fields. Valid values are: - Ignore: This will ignore any unknown fields that are silently dropped from the object, and will ignore all but the last duplicate field that the decoder encounters. This is the default behavior prior to v1.23. - Warn: This will send a warning via the standard warning response header for each unknown field that is dropped from the object, and for each duplicate field that is encountered. The request will still succeed if there are no other errors, and will only persist the last of any duplicate fields. This is the default in v1.23+ - Strict: This will fail the request with a BadRequest error if any unknown fields would be dropped from the object, or if any duplicate fields are present. The error returned from the server will contain all unknown and duplicate fields encountered. (optional) [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: tuple(object, status_code(int), headers(HTTPHeaderDict)) [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] local_var_params = locals() [e2e-llm-inference-service] [e2e-llm-inference-service] all_params = [ [e2e-llm-inference-service] 'group', [e2e-llm-inference-service] 'version', [e2e-llm-inference-service] 'namespace', [e2e-llm-inference-service] 'plural', [e2e-llm-inference-service] 'body', [e2e-llm-inference-service] 'pretty', [e2e-llm-inference-service] 'dry_run', [e2e-llm-inference-service] 'field_manager', [e2e-llm-inference-service] 'field_validation' [e2e-llm-inference-service] ] [e2e-llm-inference-service] all_params.extend( [e2e-llm-inference-service] [ [e2e-llm-inference-service] 'async_req', [e2e-llm-inference-service] '_return_http_data_only', [e2e-llm-inference-service] '_preload_content', [e2e-llm-inference-service] '_request_timeout' [e2e-llm-inference-service] ] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] for key, val in six.iteritems(local_var_params['kwargs']): [e2e-llm-inference-service] if key not in all_params: [e2e-llm-inference-service] raise ApiTypeError( [e2e-llm-inference-service] "Got an unexpected keyword argument '%s'" [e2e-llm-inference-service] " to method create_namespaced_custom_object" % key [e2e-llm-inference-service] ) [e2e-llm-inference-service] local_var_params[key] = val [e2e-llm-inference-service] del local_var_params['kwargs'] [e2e-llm-inference-service] # verify the required parameter 'group' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('group' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['group'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `group` when calling `create_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'version' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('version' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['version'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `version` when calling `create_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'namespace' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('namespace' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['namespace'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `namespace` when calling `create_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'plural' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('plural' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['plural'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `plural` when calling `create_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'body' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('body' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['body'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `body` when calling `create_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] collection_formats = {} [e2e-llm-inference-service] [e2e-llm-inference-service] path_params = {} [e2e-llm-inference-service] if 'group' in local_var_params: [e2e-llm-inference-service] path_params['group'] = local_var_params['group'] # noqa: E501 [e2e-llm-inference-service] if 'version' in local_var_params: [e2e-llm-inference-service] path_params['version'] = local_var_params['version'] # noqa: E501 [e2e-llm-inference-service] if 'namespace' in local_var_params: [e2e-llm-inference-service] path_params['namespace'] = local_var_params['namespace'] # noqa: E501 [e2e-llm-inference-service] if 'plural' in local_var_params: [e2e-llm-inference-service] path_params['plural'] = local_var_params['plural'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] if 'pretty' in local_var_params and local_var_params['pretty'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('pretty', local_var_params['pretty'])) # noqa: E501 [e2e-llm-inference-service] if 'dry_run' in local_var_params and local_var_params['dry_run'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('dryRun', local_var_params['dry_run'])) # noqa: E501 [e2e-llm-inference-service] if 'field_manager' in local_var_params and local_var_params['field_manager'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('fieldManager', local_var_params['field_manager'])) # noqa: E501 [e2e-llm-inference-service] if 'field_validation' in local_var_params and local_var_params['field_validation'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('fieldValidation', local_var_params['field_validation'])) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] header_params = {} [e2e-llm-inference-service] [e2e-llm-inference-service] form_params = [] [e2e-llm-inference-service] local_var_files = {} [e2e-llm-inference-service] [e2e-llm-inference-service] body_params = None [e2e-llm-inference-service] if 'body' in local_var_params: [e2e-llm-inference-service] body_params = local_var_params['body'] [e2e-llm-inference-service] # HTTP header `Accept` [e2e-llm-inference-service] header_params['Accept'] = self.api_client.select_header_accept( [e2e-llm-inference-service] ['application/json']) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] # Authentication setting [e2e-llm-inference-service] auth_settings = ['BearerToken'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.api_client.call_api( [e2e-llm-inference-service] '/apis/{group}/{version}/namespaces/{namespace}/{plural}', 'POST', [e2e-llm-inference-service] path_params, [e2e-llm-inference-service] query_params, [e2e-llm-inference-service] header_params, [e2e-llm-inference-service] body=body_params, [e2e-llm-inference-service] post_params=form_params, [e2e-llm-inference-service] files=local_var_files, [e2e-llm-inference-service] response_type='object', # noqa: E501 [e2e-llm-inference-service] auth_settings=auth_settings, [e2e-llm-inference-service] async_req=local_var_params.get('async_req'), [e2e-llm-inference-service] _return_http_data_only=local_var_params.get('_return_http_data_only'), # noqa: E501 [e2e-llm-inference-service] _preload_content=local_var_params.get('_preload_content', True), [e2e-llm-inference-service] _request_timeout=local_var_params.get('_request_timeout'), [e2e-llm-inference-service] collection_formats=collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:354: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/{group}/{version}/namespaces/{namespace}/{plural}' [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'namespace': 'kserve-ci-e2e-test', 'plural': 'llminferenceserviceconfigs', 'version': 'v1alpha1'} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...ustom-1626c793', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}} [e2e-llm-inference-service] post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], async_req = None, _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def call_api(self, resource_path, method, [e2e-llm-inference-service] path_params=None, query_params=None, header_params=None, [e2e-llm-inference-service] body=None, post_params=None, files=None, [e2e-llm-inference-service] response_type=None, auth_settings=None, async_req=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] """Makes the HTTP request (synchronous) and returns deserialized data. [e2e-llm-inference-service] [e2e-llm-inference-service] To make an async_req request, set the async_req parameter. [e2e-llm-inference-service] [e2e-llm-inference-service] :param resource_path: Path to method endpoint. [e2e-llm-inference-service] :param method: Method to call. [e2e-llm-inference-service] :param path_params: Path parameters in the url. [e2e-llm-inference-service] :param query_params: Query parameters in the url. [e2e-llm-inference-service] :param header_params: Header parameters to be [e2e-llm-inference-service] placed in the request header. [e2e-llm-inference-service] :param body: Request body. [e2e-llm-inference-service] :param post_params dict: Request post form parameters, [e2e-llm-inference-service] for `application/x-www-form-urlencoded`, `multipart/form-data`. [e2e-llm-inference-service] :param auth_settings list: Auth Settings names for the request. [e2e-llm-inference-service] :param response: Response data type. [e2e-llm-inference-service] :param files dict: key -> filename, value -> filepath, [e2e-llm-inference-service] for `multipart/form-data`. [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param collection_formats: dict of collection formats for path, query, [e2e-llm-inference-service] header, and post parameters. [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: [e2e-llm-inference-service] If async_req parameter is True, [e2e-llm-inference-service] the request will be called asynchronously. [e2e-llm-inference-service] The method will return the request thread. [e2e-llm-inference-service] If parameter async_req is False or missing, [e2e-llm-inference-service] then the method will return the response directly. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if not async_req: [e2e-llm-inference-service] > return self.__call_api(resource_path, method, [e2e-llm-inference-service] path_params, query_params, header_params, [e2e-llm-inference-service] body, post_params, files, [e2e-llm-inference-service] response_type, auth_settings, [e2e-llm-inference-service] _return_http_data_only, collection_formats, [e2e-llm-inference-service] _preload_content, _request_timeout, _host) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:348: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs' [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] path_params = [('group', 'serving.kserve.io'), ('version', 'v1alpha1'), ('namespace', 'kserve-ci-e2e-test'), ('plural', 'llminferenceserviceconfigs')] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...ustom-1626c793', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}} [e2e-llm-inference-service] post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def __call_api( [e2e-llm-inference-service] self, resource_path, method, path_params=None, [e2e-llm-inference-service] query_params=None, header_params=None, body=None, post_params=None, [e2e-llm-inference-service] files=None, response_type=None, auth_settings=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] [e2e-llm-inference-service] config = self.configuration [e2e-llm-inference-service] [e2e-llm-inference-service] # header parameters [e2e-llm-inference-service] header_params = header_params or {} [e2e-llm-inference-service] header_params.update(self.default_headers) [e2e-llm-inference-service] if self.cookie: [e2e-llm-inference-service] header_params['Cookie'] = self.cookie [e2e-llm-inference-service] if header_params: [e2e-llm-inference-service] header_params = self.sanitize_for_serialization(header_params) [e2e-llm-inference-service] header_params = dict(self.parameters_to_tuples(header_params, [e2e-llm-inference-service] collection_formats)) [e2e-llm-inference-service] [e2e-llm-inference-service] # path parameters [e2e-llm-inference-service] if path_params: [e2e-llm-inference-service] path_params = self.sanitize_for_serialization(path_params) [e2e-llm-inference-service] path_params = self.parameters_to_tuples(path_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] for k, v in path_params: [e2e-llm-inference-service] # specified safe chars, encode everything [e2e-llm-inference-service] resource_path = resource_path.replace( [e2e-llm-inference-service] '{%s}' % k, [e2e-llm-inference-service] quote(str(v), safe=config.safe_chars_for_path_param) [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # query parameters [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] query_params = self.sanitize_for_serialization(query_params) [e2e-llm-inference-service] query_params = self.parameters_to_tuples(query_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] # post parameters [e2e-llm-inference-service] if post_params or files: [e2e-llm-inference-service] post_params = post_params if post_params else [] [e2e-llm-inference-service] post_params = self.sanitize_for_serialization(post_params) [e2e-llm-inference-service] post_params = self.parameters_to_tuples(post_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] post_params.extend(self.files_parameters(files)) [e2e-llm-inference-service] [e2e-llm-inference-service] # auth setting [e2e-llm-inference-service] self.update_params_for_auth(header_params, query_params, auth_settings) [e2e-llm-inference-service] [e2e-llm-inference-service] # body [e2e-llm-inference-service] if body: [e2e-llm-inference-service] body = self.sanitize_for_serialization(body) [e2e-llm-inference-service] [e2e-llm-inference-service] # request url [e2e-llm-inference-service] if _host is None: [e2e-llm-inference-service] url = self.configuration.host + resource_path [e2e-llm-inference-service] else: [e2e-llm-inference-service] # use server/host defined in path or operation instead [e2e-llm-inference-service] url = _host + resource_path [e2e-llm-inference-service] [e2e-llm-inference-service] # perform request and return response [e2e-llm-inference-service] > response_data = self.request( [e2e-llm-inference-service] method, url, query_params=query_params, headers=header_params, [e2e-llm-inference-service] post_params=post_params, body=body, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:180: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] post_params = [] [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...ustom-1626c793', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}} [e2e-llm-inference-service] _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] post_params=None, body=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Makes the HTTP request using RESTClient.""" [e2e-llm-inference-service] if method == "GET": [e2e-llm-inference-service] return self.rest_client.GET(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif method == "HEAD": [e2e-llm-inference-service] return self.rest_client.HEAD(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif method == "OPTIONS": [e2e-llm-inference-service] return self.rest_client.OPTIONS(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout) [e2e-llm-inference-service] elif method == "POST": [e2e-llm-inference-service] > return self.rest_client.POST(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] post_params=post_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:391: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs' [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] query_params = [], post_params = [] [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...ustom-1626c793', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}} [e2e-llm-inference-service] _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def POST(self, url, headers=None, query_params=None, post_params=None, [e2e-llm-inference-service] body=None, _preload_content=True, _request_timeout=None): [e2e-llm-inference-service] > return self.request("POST", url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] post_params=post_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] body=body) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:279: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...ustom-1626c793', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}} [e2e-llm-inference-service] post_params = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] body=None, post_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Perform requests. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: http request method [e2e-llm-inference-service] :param url: http request url [e2e-llm-inference-service] :param query_params: query parameters in the url [e2e-llm-inference-service] :param headers: http request headers [e2e-llm-inference-service] :param body: request json body, for `application/json` [e2e-llm-inference-service] :param post_params: request post parameters, [e2e-llm-inference-service] `application/x-www-form-urlencoded` [e2e-llm-inference-service] and `multipart/form-data` [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] """ [e2e-llm-inference-service] method = method.upper() [e2e-llm-inference-service] assert method in ['GET', 'HEAD', 'DELETE', 'POST', 'PUT', [e2e-llm-inference-service] 'PATCH', 'OPTIONS'] [e2e-llm-inference-service] [e2e-llm-inference-service] if post_params and body: [e2e-llm-inference-service] raise ApiValueError( [e2e-llm-inference-service] "body parameter cannot be used with post_params parameter." [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] post_params = post_params or {} [e2e-llm-inference-service] headers = headers or {} [e2e-llm-inference-service] [e2e-llm-inference-service] timeout = None [e2e-llm-inference-service] if _request_timeout: [e2e-llm-inference-service] if isinstance(_request_timeout, (int, ) if six.PY3 else (int, long)): # noqa: E501,F821 [e2e-llm-inference-service] timeout = urllib3.Timeout(total=_request_timeout) [e2e-llm-inference-service] elif (isinstance(_request_timeout, tuple) and [e2e-llm-inference-service] len(_request_timeout) == 2): [e2e-llm-inference-service] timeout = urllib3.Timeout( [e2e-llm-inference-service] connect=_request_timeout[0], read=_request_timeout[1]) [e2e-llm-inference-service] [e2e-llm-inference-service] if 'Content-Type' not in headers: [e2e-llm-inference-service] headers['Content-Type'] = 'application/json' [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # For `POST`, `PUT`, `PATCH`, `OPTIONS`, `DELETE` [e2e-llm-inference-service] if method in ['POST', 'PUT', 'PATCH', 'OPTIONS', 'DELETE']: [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] url += '?' + urlencode(query_params) [e2e-llm-inference-service] if (re.search('json', headers['Content-Type'], re.IGNORECASE) or [e2e-llm-inference-service] headers['Content-Type'] == 'application/apply-patch+yaml'): [e2e-llm-inference-service] if headers['Content-Type'] == 'application/json-patch+json': [e2e-llm-inference-service] if not isinstance(body, list): [e2e-llm-inference-service] headers['Content-Type'] = \ [e2e-llm-inference-service] 'application/strategic-merge-patch+json' [e2e-llm-inference-service] request_body = None [e2e-llm-inference-service] if body is not None: [e2e-llm-inference-service] request_body = json.dumps(body) [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'application/x-www-form-urlencoded': # noqa: E501 [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=False, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'multipart/form-data': [e2e-llm-inference-service] # must del headers['Content-Type'], or the correct [e2e-llm-inference-service] # Content-Type which generated by urllib3 will be [e2e-llm-inference-service] # overwritten. [e2e-llm-inference-service] del headers['Content-Type'] [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=True, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] # Pass a `string` parameter directly in the body to support [e2e-llm-inference-service] # other content types than Json when `body` argument is [e2e-llm-inference-service] # provided in serialized form [e2e-llm-inference-service] elif isinstance(body, str) or isinstance(body, bytes): [e2e-llm-inference-service] request_body = body [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Cannot generate the request from given parameters [e2e-llm-inference-service] msg = """Cannot prepare a request message for provided [e2e-llm-inference-service] arguments. Please check that your arguments match [e2e-llm-inference-service] declared content type.""" [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] # For `GET`, `HEAD` [e2e-llm-inference-service] else: [e2e-llm-inference-service] r = self.pool_manager.request(method, url, [e2e-llm-inference-service] fields=query_params, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] except urllib3.exceptions.SSLError as e: [e2e-llm-inference-service] msg = "{0}\n{1}".format(type(e).__name__, str(e)) [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] [e2e-llm-inference-service] if _preload_content: [e2e-llm-inference-service] r = RESTResponse(r) [e2e-llm-inference-service] [e2e-llm-inference-service] # In the python 3, the response.data is bytes. [e2e-llm-inference-service] # we need to decode it to string. [e2e-llm-inference-service] if six.PY3: [e2e-llm-inference-service] r.data = r.data.decode('utf8') [e2e-llm-inference-service] [e2e-llm-inference-service] # log response body [e2e-llm-inference-service] logger.debug("response body: %s", r.data) [e2e-llm-inference-service] [e2e-llm-inference-service] if not 200 <= r.status <= 299: [e2e-llm-inference-service] > raise ApiException(http_resp=r) [e2e-llm-inference-service] E kubernetes.client.exceptions.ApiException: (500) [e2e-llm-inference-service] E Reason: Internal Server Error [e2e-llm-inference-service] E HTTP response headers: HTTPHeaderDict({'Audit-Id': 'cc66f3d7-4dc5-4139-9c7e-c390de72bf14', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': 'a8664215-74a2-48d8-b231-e657e37b3200', 'X-Kubernetes-Pf-Prioritylevel-Uid': '0e7454c5-cfcf-4e4f-b47e-e88e460e3769', 'Date': 'Thu, 02 Jul 2026 14:29:47 GMT', 'Content-Length': '701'}) [e2e-llm-inference-service] E HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"Internal error occurred: failed calling webhook \"llminferenceserviceconfig.kserve-webhook-server.v1alpha1.validator\": failed to call webhook: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/validate-serving-kserve-io-v1alpha1-llminferenceserviceconfig?timeout=10s\": EOF","reason":"InternalError","details":{"causes":[{"message":"failed calling webhook \"llminferenceserviceconfig.kserve-webhook-server.v1alpha1.validator\": failed to call webhook: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/validate-serving-kserve-io-v1alpha1-llminferenceserviceconfig?timeout=10s\": EOF"}]},"code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:238: ApiException [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig router-managed-scheduler-custom-1626c793 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig router-managed-scheduler-custom-1626c793 [e2e-llm-inference-service] _ ERROR at setup of test_rolling_upgrade_coordination[router-managed-workload-llmd-simulator-model-fb-opt-125m] _ [e2e-llm-inference-service] [gw1] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def _new_conn(self) -> socket.socket: [e2e-llm-inference-service] """Establish a socket connection and set nodelay settings on it. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: New socket connection. [e2e-llm-inference-service] """ [e2e-llm-inference-service] try: [e2e-llm-inference-service] > sock = connection.create_connection( [e2e-llm-inference-service] (self._dns_host, self.port), [e2e-llm-inference-service] self.timeout, [e2e-llm-inference-service] source_address=self.source_address, [e2e-llm-inference-service] socket_options=self.socket_options, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:204: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] address = ('a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', 6443) [e2e-llm-inference-service] timeout = None, source_address = None, socket_options = [(6, 1, 1)] [e2e-llm-inference-service] [e2e-llm-inference-service] def create_connection( [e2e-llm-inference-service] address: tuple[str, int], [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] source_address: tuple[str, int] | None = None, [e2e-llm-inference-service] socket_options: _TYPE_SOCKET_OPTIONS | None = None, [e2e-llm-inference-service] ) -> socket.socket: [e2e-llm-inference-service] """Connect to *address* and return the socket object. [e2e-llm-inference-service] [e2e-llm-inference-service] Convenience function. Connect to *address* (a 2-tuple ``(host, [e2e-llm-inference-service] port)``) and return the socket object. Passing the optional [e2e-llm-inference-service] *timeout* parameter will set the timeout on the socket instance [e2e-llm-inference-service] before attempting to connect. If no *timeout* is supplied, the [e2e-llm-inference-service] global default timeout setting returned by :func:`socket.getdefaulttimeout` [e2e-llm-inference-service] is used. If *source_address* is set it must be a tuple of (host, port) [e2e-llm-inference-service] for the socket to bind as a source address before making the connection. [e2e-llm-inference-service] An host of '' or port 0 tells the OS to use the default. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] host, port = address [e2e-llm-inference-service] if host.startswith("["): [e2e-llm-inference-service] host = host.strip("[]") [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Using the value from allowed_gai_family() in the context of getaddrinfo lets [e2e-llm-inference-service] # us select whether to work with IPv4 DNS records, IPv6 records, or both. [e2e-llm-inference-service] # The original create_connection function always returns all records. [e2e-llm-inference-service] family = allowed_gai_family() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] host.encode("idna") [e2e-llm-inference-service] except UnicodeError: [e2e-llm-inference-service] raise LocationParseError(f"'{host}', label empty or too long") from None [e2e-llm-inference-service] [e2e-llm-inference-service] > for res in socket.getaddrinfo(host, port, family, socket.SOCK_STREAM): [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/util/connection.py:60: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] host = 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' [e2e-llm-inference-service] port = 6443, family = [e2e-llm-inference-service] type = , proto = 0, flags = 0 [e2e-llm-inference-service] [e2e-llm-inference-service] def getaddrinfo(host, port, family=0, type=0, proto=0, flags=0): [e2e-llm-inference-service] """Resolve host and port into list of address info entries. [e2e-llm-inference-service] [e2e-llm-inference-service] Translate the host/port argument into a sequence of 5-tuples that contain [e2e-llm-inference-service] all the necessary arguments for creating a socket connected to that service. [e2e-llm-inference-service] host is a domain name, a string representation of an IPv4/v6 address or [e2e-llm-inference-service] None. port is a string service name such as 'http', a numeric port number or [e2e-llm-inference-service] None. By passing None as the value of host and port, you can pass NULL to [e2e-llm-inference-service] the underlying C API. [e2e-llm-inference-service] [e2e-llm-inference-service] The family, type and proto arguments can be optionally specified in order to [e2e-llm-inference-service] narrow the list of addresses returned. Passing zero as a value for each of [e2e-llm-inference-service] these arguments selects the full range of results. [e2e-llm-inference-service] """ [e2e-llm-inference-service] # We override this function since we want to translate the numeric family [e2e-llm-inference-service] # and socket type values to enum constants. [e2e-llm-inference-service] addrlist = [] [e2e-llm-inference-service] > for res in _socket.getaddrinfo(host, port, family, type, proto, flags): [e2e-llm-inference-service] E socket.gaierror: [Errno -2] Name or service not known [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/socket.py:974: gaierror [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False, timeout = None, pool_timeout = None [e2e-llm-inference-service] release_conn = True, chunked = False, body_pos = None, preload_content = True [e2e-llm-inference-service] decode_content = True, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] > response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:787: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=None, read=None, total=None), chunked = False [e2e-llm-inference-service] response_conn = None, preload_content = True, decode_content = True [e2e-llm-inference-service] enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] > raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:488: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=None, read=None, total=None), chunked = False [e2e-llm-inference-service] response_conn = None, preload_content = True, decode_content = True [e2e-llm-inference-service] enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] > self._validate_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:464: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] [e2e-llm-inference-service] def _validate_conn(self, conn: BaseHTTPConnection) -> None: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Called right before a request is made, after the socket is created. [e2e-llm-inference-service] """ [e2e-llm-inference-service] super()._validate_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] # Force connect early to allow us to validate the connection. [e2e-llm-inference-service] if conn.is_closed: [e2e-llm-inference-service] > conn.connect() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:1093: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def connect(self) -> None: [e2e-llm-inference-service] # Today we don't need to be doing this step before the /actual/ socket [e2e-llm-inference-service] # connection, however in the future we'll need to decide whether to [e2e-llm-inference-service] # create a new socket or re-use an existing "shared" socket as a part [e2e-llm-inference-service] # of the HTTP/2 handshake dance. [e2e-llm-inference-service] if self._tunnel_host is not None and self._tunnel_port is not None: [e2e-llm-inference-service] probe_http2_host = self._tunnel_host [e2e-llm-inference-service] probe_http2_port = self._tunnel_port [e2e-llm-inference-service] else: [e2e-llm-inference-service] probe_http2_host = self.host [e2e-llm-inference-service] probe_http2_port = self.port [e2e-llm-inference-service] [e2e-llm-inference-service] # Check if the target origin supports HTTP/2. [e2e-llm-inference-service] # If the value comes back as 'None' it means that the current thread [e2e-llm-inference-service] # is probing for HTTP/2 support. Otherwise, we're waiting for another [e2e-llm-inference-service] # probe to complete, or we get a value right away. [e2e-llm-inference-service] target_supports_http2: bool | None [e2e-llm-inference-service] if "h2" in ssl_.ALPN_PROTOCOLS: [e2e-llm-inference-service] target_supports_http2 = http2_probe.acquire_and_get( [e2e-llm-inference-service] host=probe_http2_host, port=probe_http2_port [e2e-llm-inference-service] ) [e2e-llm-inference-service] else: [e2e-llm-inference-service] # If HTTP/2 isn't going to be offered it doesn't matter if [e2e-llm-inference-service] # the target supports HTTP/2. Don't want to make a probe. [e2e-llm-inference-service] target_supports_http2 = False [e2e-llm-inference-service] [e2e-llm-inference-service] if self._connect_callback is not None: [e2e-llm-inference-service] self._connect_callback( [e2e-llm-inference-service] "before connect", [e2e-llm-inference-service] thread_id=threading.get_ident(), [e2e-llm-inference-service] target_supports_http2=target_supports_http2, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] sock: socket.socket | ssl.SSLSocket [e2e-llm-inference-service] > self.sock = sock = self._new_conn() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:759: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def _new_conn(self) -> socket.socket: [e2e-llm-inference-service] """Establish a socket connection and set nodelay settings on it. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: New socket connection. [e2e-llm-inference-service] """ [e2e-llm-inference-service] try: [e2e-llm-inference-service] sock = connection.create_connection( [e2e-llm-inference-service] (self._dns_host, self.port), [e2e-llm-inference-service] self.timeout, [e2e-llm-inference-service] source_address=self.source_address, [e2e-llm-inference-service] socket_options=self.socket_options, [e2e-llm-inference-service] ) [e2e-llm-inference-service] except socket.gaierror as e: [e2e-llm-inference-service] > raise NameResolutionError(self.host, self, e) from e [e2e-llm-inference-service] E urllib3.exceptions.NameResolutionError: HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Failed to resolve 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:211: NameResolutionError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] request = > [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.fixture(scope="function") [e2e-llm-inference-service] def test_case(request): [e2e-llm-inference-service] tc = request.param [e2e-llm-inference-service] [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Execute before test hooks [e2e-llm-inference-service] try: [e2e-llm-inference-service] for func in tc.before_test: [e2e-llm-inference-service] func() [e2e-llm-inference-service] except Exception as before_test_error: [e2e-llm-inference-service] raise RuntimeError( [e2e-llm-inference-service] f"Failed to execute before test hook: {before_test_error}" [e2e-llm-inference-service] ) from before_test_error [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > _setup_test_case_service(kserve_client, tc, request.node.name) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1476: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] tc = TestCase(base_refs=['router-managed', 'workload-llmd-simulator', 'model-fb-opt-125m'], prompt='KServe is a', service_n...None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service=None, model_name='facebook/opt-125m') [e2e-llm-inference-service] test_node_name = 'test_rolling_upgrade_coordination[router-managed-workload-llmd-simulator-model-fb-opt-125m]' [e2e-llm-inference-service] peer_index = None [e2e-llm-inference-service] [e2e-llm-inference-service] def _setup_test_case_service(kserve_client, tc, test_node_name, peer_index=None): [e2e-llm-inference-service] """Create LLMInferenceServiceConfigs and build the LLMInferenceService for a TestCase. [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a list of created config names for cleanup tracking. [e2e-llm-inference-service] """ [e2e-llm-inference-service] missing_refs = [ [e2e-llm-inference-service] ref for ref in tc.base_refs if ref not in LLMINFERENCESERVICE_CONFIGS [e2e-llm-inference-service] ] [e2e-llm-inference-service] if missing_refs: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Missing base_refs in LLMINFERENCESERVICE_CONFIGS: {missing_refs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] if not tc.service_name: [e2e-llm-inference-service] suffix = f"-peer-{peer_index}" if peer_index is not None else "" [e2e-llm-inference-service] tc.service_name = generate_service_name(test_node_name + suffix, tc.base_refs) [e2e-llm-inference-service] if tc.model_name == "default/model": [e2e-llm-inference-service] tc.model_name = _get_model_name_from_configs(tc.base_refs) [e2e-llm-inference-service] [e2e-llm-inference-service] created_configs = [] [e2e-llm-inference-service] unique_base_refs = [] [e2e-llm-inference-service] for base_ref in tc.base_refs: [e2e-llm-inference-service] unique_config_name = generate_k8s_safe_suffix(base_ref, [tc.service_name]) [e2e-llm-inference-service] unique_base_refs.append(unique_config_name) [e2e-llm-inference-service] [e2e-llm-inference-service] unique_config_body = { [e2e-llm-inference-service] "apiVersion": "serving.kserve.io/v1alpha1", [e2e-llm-inference-service] "kind": "LLMInferenceServiceConfig", [e2e-llm-inference-service] "metadata": { [e2e-llm-inference-service] "name": unique_config_name, [e2e-llm-inference-service] "namespace": KSERVE_TEST_NAMESPACE, [e2e-llm-inference-service] }, [e2e-llm-inference-service] "spec": LLMINFERENCESERVICE_CONFIGS[base_ref], [e2e-llm-inference-service] } [e2e-llm-inference-service] [e2e-llm-inference-service] > _create_or_update_llmisvc_config( [e2e-llm-inference-service] kserve_client, unique_config_body, KSERVE_TEST_NAMESPACE [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1436: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] llm_config = {'apiVersion': 'serving.kserve.io/v1alpha1', 'kind': 'LLMInferenceServiceConfig', 'metadata': {'name': 'router-managed...grade-64e041bd', 'namespace': 'kserve-ci-e2e-test'}, 'spec': {'router': {'gateway': {}, 'route': {}, 'scheduler': {}}}} [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test' [e2e-llm-inference-service] [e2e-llm-inference-service] def _create_or_update_llmisvc_config(kserve_client, llm_config, namespace=None): [e2e-llm-inference-service] """Create or update an LLMInferenceServiceConfig resource.""" [e2e-llm-inference-service] version = llm_config["apiVersion"].split("/")[1] [e2e-llm-inference-service] [e2e-llm-inference-service] if namespace is None: [e2e-llm-inference-service] namespace = llm_config.get("metadata", {}).get("namespace", "default") [e2e-llm-inference-service] [e2e-llm-inference-service] name = llm_config.get("metadata", {}).get("name") [e2e-llm-inference-service] if not name: [e2e-llm-inference-service] raise ValueError("LLMInferenceServiceConfig must have a name in metadata") [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Checking LLMInferenceServiceConfig {name} in namespace {namespace}") [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > existing_config = kserve_client.api_instance.get_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICECONFIG, [e2e-llm-inference-service] name, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/fixtures.py:1589: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test', plural = 'llminferenceserviceconfigs' [e2e-llm-inference-service] name = 'router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] [e2e-llm-inference-service] def get_namespaced_custom_object(self, group, version, namespace, plural, name, **kwargs): # noqa: E501 [e2e-llm-inference-service] """get_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a namespace scoped custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.get_namespaced_custom_object(group, version, namespace, plural, name, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: the custom resource's group (required) [e2e-llm-inference-service] :param str version: the custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: the custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param str name: the custom object's name (required) [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: object [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] kwargs['_return_http_data_only'] = True [e2e-llm-inference-service] > return self.get_namespaced_custom_object_with_http_info(group, version, namespace, plural, name, **kwargs) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:1632: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test', plural = 'llminferenceserviceconfigs' [e2e-llm-inference-service] name = 'router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] local_var_params = {'_return_http_data_only': True, 'all_params': ['group', 'version', 'namespace', 'plural', 'name', 'async_req', ...], 'auth_settings': ['BearerToken'], 'body_params': None, ...} [e2e-llm-inference-service] all_params = ['group', 'version', 'namespace', 'plural', 'name', 'async_req', ...] [e2e-llm-inference-service] key = '_return_http_data_only', val = True, collection_formats = {} [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'name': 'router-managed-rolling-upgrade-64e041bd', 'namespace': 'kserve-ci-e2e-test', 'plural': 'llminferenceserviceconfigs', ...} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] def get_namespaced_custom_object_with_http_info(self, group, version, namespace, plural, name, **kwargs): # noqa: E501 [e2e-llm-inference-service] """get_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a namespace scoped custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.get_namespaced_custom_object_with_http_info(group, version, namespace, plural, name, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: the custom resource's group (required) [e2e-llm-inference-service] :param str version: the custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: the custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param str name: the custom object's name (required) [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: tuple(object, status_code(int), headers(HTTPHeaderDict)) [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] local_var_params = locals() [e2e-llm-inference-service] [e2e-llm-inference-service] all_params = [ [e2e-llm-inference-service] 'group', [e2e-llm-inference-service] 'version', [e2e-llm-inference-service] 'namespace', [e2e-llm-inference-service] 'plural', [e2e-llm-inference-service] 'name' [e2e-llm-inference-service] ] [e2e-llm-inference-service] all_params.extend( [e2e-llm-inference-service] [ [e2e-llm-inference-service] 'async_req', [e2e-llm-inference-service] '_return_http_data_only', [e2e-llm-inference-service] '_preload_content', [e2e-llm-inference-service] '_request_timeout' [e2e-llm-inference-service] ] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] for key, val in six.iteritems(local_var_params['kwargs']): [e2e-llm-inference-service] if key not in all_params: [e2e-llm-inference-service] raise ApiTypeError( [e2e-llm-inference-service] "Got an unexpected keyword argument '%s'" [e2e-llm-inference-service] " to method get_namespaced_custom_object" % key [e2e-llm-inference-service] ) [e2e-llm-inference-service] local_var_params[key] = val [e2e-llm-inference-service] del local_var_params['kwargs'] [e2e-llm-inference-service] # verify the required parameter 'group' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('group' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['group'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `group` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'version' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('version' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['version'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `version` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'namespace' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('namespace' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['namespace'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `namespace` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'plural' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('plural' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['plural'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `plural` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'name' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('name' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['name'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `name` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] collection_formats = {} [e2e-llm-inference-service] [e2e-llm-inference-service] path_params = {} [e2e-llm-inference-service] if 'group' in local_var_params: [e2e-llm-inference-service] path_params['group'] = local_var_params['group'] # noqa: E501 [e2e-llm-inference-service] if 'version' in local_var_params: [e2e-llm-inference-service] path_params['version'] = local_var_params['version'] # noqa: E501 [e2e-llm-inference-service] if 'namespace' in local_var_params: [e2e-llm-inference-service] path_params['namespace'] = local_var_params['namespace'] # noqa: E501 [e2e-llm-inference-service] if 'plural' in local_var_params: [e2e-llm-inference-service] path_params['plural'] = local_var_params['plural'] # noqa: E501 [e2e-llm-inference-service] if 'name' in local_var_params: [e2e-llm-inference-service] path_params['name'] = local_var_params['name'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] header_params = {} [e2e-llm-inference-service] [e2e-llm-inference-service] form_params = [] [e2e-llm-inference-service] local_var_files = {} [e2e-llm-inference-service] [e2e-llm-inference-service] body_params = None [e2e-llm-inference-service] # HTTP header `Accept` [e2e-llm-inference-service] header_params['Accept'] = self.api_client.select_header_accept( [e2e-llm-inference-service] ['application/json']) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] # Authentication setting [e2e-llm-inference-service] auth_settings = ['BearerToken'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.api_client.call_api( [e2e-llm-inference-service] '/apis/{group}/{version}/namespaces/{namespace}/{plural}/{name}', 'GET', [e2e-llm-inference-service] path_params, [e2e-llm-inference-service] query_params, [e2e-llm-inference-service] header_params, [e2e-llm-inference-service] body=body_params, [e2e-llm-inference-service] post_params=form_params, [e2e-llm-inference-service] files=local_var_files, [e2e-llm-inference-service] response_type='object', # noqa: E501 [e2e-llm-inference-service] auth_settings=auth_settings, [e2e-llm-inference-service] async_req=local_var_params.get('async_req'), [e2e-llm-inference-service] _return_http_data_only=local_var_params.get('_return_http_data_only'), # noqa: E501 [e2e-llm-inference-service] _preload_content=local_var_params.get('_preload_content', True), [e2e-llm-inference-service] _request_timeout=local_var_params.get('_request_timeout'), [e2e-llm-inference-service] collection_formats=collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:1739: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/{group}/{version}/namespaces/{namespace}/{plural}/{name}' [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'name': 'router-managed-rolling-upgrade-64e041bd', 'namespace': 'kserve-ci-e2e-test', 'plural': 'llminferenceserviceconfigs', ...} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], async_req = None, _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def call_api(self, resource_path, method, [e2e-llm-inference-service] path_params=None, query_params=None, header_params=None, [e2e-llm-inference-service] body=None, post_params=None, files=None, [e2e-llm-inference-service] response_type=None, auth_settings=None, async_req=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] """Makes the HTTP request (synchronous) and returns deserialized data. [e2e-llm-inference-service] [e2e-llm-inference-service] To make an async_req request, set the async_req parameter. [e2e-llm-inference-service] [e2e-llm-inference-service] :param resource_path: Path to method endpoint. [e2e-llm-inference-service] :param method: Method to call. [e2e-llm-inference-service] :param path_params: Path parameters in the url. [e2e-llm-inference-service] :param query_params: Query parameters in the url. [e2e-llm-inference-service] :param header_params: Header parameters to be [e2e-llm-inference-service] placed in the request header. [e2e-llm-inference-service] :param body: Request body. [e2e-llm-inference-service] :param post_params dict: Request post form parameters, [e2e-llm-inference-service] for `application/x-www-form-urlencoded`, `multipart/form-data`. [e2e-llm-inference-service] :param auth_settings list: Auth Settings names for the request. [e2e-llm-inference-service] :param response: Response data type. [e2e-llm-inference-service] :param files dict: key -> filename, value -> filepath, [e2e-llm-inference-service] for `multipart/form-data`. [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param collection_formats: dict of collection formats for path, query, [e2e-llm-inference-service] header, and post parameters. [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: [e2e-llm-inference-service] If async_req parameter is True, [e2e-llm-inference-service] the request will be called asynchronously. [e2e-llm-inference-service] The method will return the request thread. [e2e-llm-inference-service] If parameter async_req is False or missing, [e2e-llm-inference-service] then the method will return the response directly. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if not async_req: [e2e-llm-inference-service] > return self.__call_api(resource_path, method, [e2e-llm-inference-service] path_params, query_params, header_params, [e2e-llm-inference-service] body, post_params, files, [e2e-llm-inference-service] response_type, auth_settings, [e2e-llm-inference-service] _return_http_data_only, collection_formats, [e2e-llm-inference-service] _preload_content, _request_timeout, _host) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:348: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] path_params = [('group', 'serving.kserve.io'), ('version', 'v1alpha1'), ('namespace', 'kserve-ci-e2e-test'), ('plural', 'llminferenceserviceconfigs'), ('name', 'router-managed-rolling-upgrade-64e041bd')] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def __call_api( [e2e-llm-inference-service] self, resource_path, method, path_params=None, [e2e-llm-inference-service] query_params=None, header_params=None, body=None, post_params=None, [e2e-llm-inference-service] files=None, response_type=None, auth_settings=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] [e2e-llm-inference-service] config = self.configuration [e2e-llm-inference-service] [e2e-llm-inference-service] # header parameters [e2e-llm-inference-service] header_params = header_params or {} [e2e-llm-inference-service] header_params.update(self.default_headers) [e2e-llm-inference-service] if self.cookie: [e2e-llm-inference-service] header_params['Cookie'] = self.cookie [e2e-llm-inference-service] if header_params: [e2e-llm-inference-service] header_params = self.sanitize_for_serialization(header_params) [e2e-llm-inference-service] header_params = dict(self.parameters_to_tuples(header_params, [e2e-llm-inference-service] collection_formats)) [e2e-llm-inference-service] [e2e-llm-inference-service] # path parameters [e2e-llm-inference-service] if path_params: [e2e-llm-inference-service] path_params = self.sanitize_for_serialization(path_params) [e2e-llm-inference-service] path_params = self.parameters_to_tuples(path_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] for k, v in path_params: [e2e-llm-inference-service] # specified safe chars, encode everything [e2e-llm-inference-service] resource_path = resource_path.replace( [e2e-llm-inference-service] '{%s}' % k, [e2e-llm-inference-service] quote(str(v), safe=config.safe_chars_for_path_param) [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # query parameters [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] query_params = self.sanitize_for_serialization(query_params) [e2e-llm-inference-service] query_params = self.parameters_to_tuples(query_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] # post parameters [e2e-llm-inference-service] if post_params or files: [e2e-llm-inference-service] post_params = post_params if post_params else [] [e2e-llm-inference-service] post_params = self.sanitize_for_serialization(post_params) [e2e-llm-inference-service] post_params = self.parameters_to_tuples(post_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] post_params.extend(self.files_parameters(files)) [e2e-llm-inference-service] [e2e-llm-inference-service] # auth setting [e2e-llm-inference-service] self.update_params_for_auth(header_params, query_params, auth_settings) [e2e-llm-inference-service] [e2e-llm-inference-service] # body [e2e-llm-inference-service] if body: [e2e-llm-inference-service] body = self.sanitize_for_serialization(body) [e2e-llm-inference-service] [e2e-llm-inference-service] # request url [e2e-llm-inference-service] if _host is None: [e2e-llm-inference-service] url = self.configuration.host + resource_path [e2e-llm-inference-service] else: [e2e-llm-inference-service] # use server/host defined in path or operation instead [e2e-llm-inference-service] url = _host + resource_path [e2e-llm-inference-service] [e2e-llm-inference-service] # perform request and return response [e2e-llm-inference-service] > response_data = self.request( [e2e-llm-inference-service] method, url, query_params=query_params, headers=header_params, [e2e-llm-inference-service] post_params=post_params, body=body, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:180: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] post_params = [], body = None, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] post_params=None, body=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Makes the HTTP request using RESTClient.""" [e2e-llm-inference-service] if method == "GET": [e2e-llm-inference-service] > return self.rest_client.GET(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:373: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] query_params = [], _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def GET(self, url, headers=None, query_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] > return self.request("GET", url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] query_params=query_params) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:244: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] body=None, post_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Perform requests. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: http request method [e2e-llm-inference-service] :param url: http request url [e2e-llm-inference-service] :param query_params: query parameters in the url [e2e-llm-inference-service] :param headers: http request headers [e2e-llm-inference-service] :param body: request json body, for `application/json` [e2e-llm-inference-service] :param post_params: request post parameters, [e2e-llm-inference-service] `application/x-www-form-urlencoded` [e2e-llm-inference-service] and `multipart/form-data` [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] """ [e2e-llm-inference-service] method = method.upper() [e2e-llm-inference-service] assert method in ['GET', 'HEAD', 'DELETE', 'POST', 'PUT', [e2e-llm-inference-service] 'PATCH', 'OPTIONS'] [e2e-llm-inference-service] [e2e-llm-inference-service] if post_params and body: [e2e-llm-inference-service] raise ApiValueError( [e2e-llm-inference-service] "body parameter cannot be used with post_params parameter." [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] post_params = post_params or {} [e2e-llm-inference-service] headers = headers or {} [e2e-llm-inference-service] [e2e-llm-inference-service] timeout = None [e2e-llm-inference-service] if _request_timeout: [e2e-llm-inference-service] if isinstance(_request_timeout, (int, ) if six.PY3 else (int, long)): # noqa: E501,F821 [e2e-llm-inference-service] timeout = urllib3.Timeout(total=_request_timeout) [e2e-llm-inference-service] elif (isinstance(_request_timeout, tuple) and [e2e-llm-inference-service] len(_request_timeout) == 2): [e2e-llm-inference-service] timeout = urllib3.Timeout( [e2e-llm-inference-service] connect=_request_timeout[0], read=_request_timeout[1]) [e2e-llm-inference-service] [e2e-llm-inference-service] if 'Content-Type' not in headers: [e2e-llm-inference-service] headers['Content-Type'] = 'application/json' [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # For `POST`, `PUT`, `PATCH`, `OPTIONS`, `DELETE` [e2e-llm-inference-service] if method in ['POST', 'PUT', 'PATCH', 'OPTIONS', 'DELETE']: [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] url += '?' + urlencode(query_params) [e2e-llm-inference-service] if (re.search('json', headers['Content-Type'], re.IGNORECASE) or [e2e-llm-inference-service] headers['Content-Type'] == 'application/apply-patch+yaml'): [e2e-llm-inference-service] if headers['Content-Type'] == 'application/json-patch+json': [e2e-llm-inference-service] if not isinstance(body, list): [e2e-llm-inference-service] headers['Content-Type'] = \ [e2e-llm-inference-service] 'application/strategic-merge-patch+json' [e2e-llm-inference-service] request_body = None [e2e-llm-inference-service] if body is not None: [e2e-llm-inference-service] request_body = json.dumps(body) [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'application/x-www-form-urlencoded': # noqa: E501 [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=False, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'multipart/form-data': [e2e-llm-inference-service] # must del headers['Content-Type'], or the correct [e2e-llm-inference-service] # Content-Type which generated by urllib3 will be [e2e-llm-inference-service] # overwritten. [e2e-llm-inference-service] del headers['Content-Type'] [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=True, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] # Pass a `string` parameter directly in the body to support [e2e-llm-inference-service] # other content types than Json when `body` argument is [e2e-llm-inference-service] # provided in serialized form [e2e-llm-inference-service] elif isinstance(body, str) or isinstance(body, bytes): [e2e-llm-inference-service] request_body = body [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Cannot generate the request from given parameters [e2e-llm-inference-service] msg = """Cannot prepare a request message for provided [e2e-llm-inference-service] arguments. Please check that your arguments match [e2e-llm-inference-service] declared content type.""" [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] # For `GET`, `HEAD` [e2e-llm-inference-service] else: [e2e-llm-inference-service] > r = self.pool_manager.request(method, url, [e2e-llm-inference-service] fields=query_params, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:217: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] body = None, fields = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] json = None, urlopen_kw = {'preload_content': True, 'timeout': None} [e2e-llm-inference-service] [e2e-llm-inference-service] def request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] fields: _TYPE_FIELDS | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] json: typing.Any | None = None, [e2e-llm-inference-service] **urlopen_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Make a request using :meth:`urlopen` with the appropriate encoding of [e2e-llm-inference-service] ``fields`` based on the ``method`` used. [e2e-llm-inference-service] [e2e-llm-inference-service] This is a convenience method that requires the least amount of manual [e2e-llm-inference-service] effort. It can be used in most situations, while still having the [e2e-llm-inference-service] option to drop down to more specific methods when necessary, such as [e2e-llm-inference-service] :meth:`request_encode_url`, :meth:`request_encode_body`, [e2e-llm-inference-service] or even the lowest level :meth:`urlopen`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param fields: [e2e-llm-inference-service] Data to encode and send in the URL or request body, depending on ``method``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param json: [e2e-llm-inference-service] Data to encode and send as JSON with UTF-encoded in the request body. [e2e-llm-inference-service] The ``"Content-Type"`` header will be set to ``"application/json"`` [e2e-llm-inference-service] unless specified otherwise. [e2e-llm-inference-service] """ [e2e-llm-inference-service] method = method.upper() [e2e-llm-inference-service] [e2e-llm-inference-service] if json is not None and body is not None: [e2e-llm-inference-service] raise TypeError( [e2e-llm-inference-service] "request got values for both 'body' and 'json' parameters which are mutually exclusive" [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if json is not None: [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not ("content-type" in map(str.lower, headers.keys())): [e2e-llm-inference-service] headers = HTTPHeaderDict(headers) [e2e-llm-inference-service] headers["Content-Type"] = "application/json" [e2e-llm-inference-service] [e2e-llm-inference-service] body = _json.dumps(json, separators=(",", ":"), ensure_ascii=False).encode( [e2e-llm-inference-service] "utf-8" [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if body is not None: [e2e-llm-inference-service] urlopen_kw["body"] = body [e2e-llm-inference-service] [e2e-llm-inference-service] if method in self._encode_url_methods: [e2e-llm-inference-service] > return self.request_encode_url( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] fields=fields, # type: ignore[arg-type] [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] **urlopen_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/_request_methods.py:135: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] fields = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] urlopen_kw = {'preload_content': True, 'timeout': None} [e2e-llm-inference-service] extra_kw = {'headers': {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'}, 'preload_content': True, 'timeout': None} [e2e-llm-inference-service] [e2e-llm-inference-service] def request_encode_url( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] fields: _TYPE_ENCODE_URL_FIELDS | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] **urlopen_kw: str, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Make a request using :meth:`urlopen` with the ``fields`` encoded in [e2e-llm-inference-service] the url. This is useful for request methods like GET, HEAD, DELETE, etc. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param fields: [e2e-llm-inference-service] Data to encode and send in the URL. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] extra_kw: dict[str, typing.Any] = {"headers": headers} [e2e-llm-inference-service] extra_kw.update(urlopen_kw) [e2e-llm-inference-service] [e2e-llm-inference-service] if fields: [e2e-llm-inference-service] url += "?" + urlencode(fields) [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.urlopen(method, url, **extra_kw) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/_request_methods.py:182: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] redirect = True [e2e-llm-inference-service] kw = {'assert_same_host': False, 'headers': {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'}, 'preload_content': True, 'redirect': False, ...} [e2e-llm-inference-service] u = Url(scheme='https', auth=None, host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', p...aces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd', query=None, fragment=None) [e2e-llm-inference-service] conn = [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, method: str, url: str, redirect: bool = True, **kw: typing.Any [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Same as :meth:`urllib3.HTTPConnectionPool.urlopen` [e2e-llm-inference-service] with custom cross-host redirect logic and only sends the request-uri [e2e-llm-inference-service] portion of the ``url``. [e2e-llm-inference-service] [e2e-llm-inference-service] The given ``url`` parameter must be absolute, such that an appropriate [e2e-llm-inference-service] :class:`urllib3.connectionpool.ConnectionPool` can be chosen for it. [e2e-llm-inference-service] """ [e2e-llm-inference-service] u = parse_url(url) [e2e-llm-inference-service] [e2e-llm-inference-service] if u.scheme is None: [e2e-llm-inference-service] warnings.warn( [e2e-llm-inference-service] "URLs without a scheme (ie 'https://') are deprecated and will raise an error " [e2e-llm-inference-service] "in a future version of urllib3. To avoid this DeprecationWarning ensure all URLs " [e2e-llm-inference-service] "start with 'https://' or 'http://'. Read more in this issue: " [e2e-llm-inference-service] "https://github.com/urllib3/urllib3/issues/2920", [e2e-llm-inference-service] category=DeprecationWarning, [e2e-llm-inference-service] stacklevel=2, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = self.connection_from_host(u.host, port=u.port, scheme=u.scheme) [e2e-llm-inference-service] [e2e-llm-inference-service] kw["assert_same_host"] = False [e2e-llm-inference-service] kw["redirect"] = False [e2e-llm-inference-service] [e2e-llm-inference-service] if "headers" not in kw: [e2e-llm-inference-service] kw["headers"] = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if self._proxy_requires_url_absolute_form(u): [e2e-llm-inference-service] response = conn.urlopen(method, url, **kw) [e2e-llm-inference-service] else: [e2e-llm-inference-service] > response = conn.urlopen(method, u.request_uri, **kw) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/poolmanager.py:457: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=2, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False, timeout = None, pool_timeout = None [e2e-llm-inference-service] release_conn = True, chunked = False, body_pos = None, preload_content = True [e2e-llm-inference-service] decode_content = True, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.c...a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=1, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False, timeout = None, pool_timeout = None [e2e-llm-inference-service] release_conn = True, chunked = False, body_pos = None, preload_content = True [e2e-llm-inference-service] decode_content = True, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.c...a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False, timeout = None, pool_timeout = None [e2e-llm-inference-service] release_conn = True, chunked = False, body_pos = None, preload_content = True [e2e-llm-inference-service] decode_content = True, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.c...a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False, timeout = None, pool_timeout = None [e2e-llm-inference-service] release_conn = True, chunked = False, body_pos = None, preload_content = True [e2e-llm-inference-service] decode_content = True, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] > retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:841: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd' [e2e-llm-inference-service] response = None [e2e-llm-inference-service] error = NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.c...a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)") [e2e-llm-inference-service] _pool = [e2e-llm-inference-service] _stacktrace = [e2e-llm-inference-service] [e2e-llm-inference-service] def increment( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str | None = None, [e2e-llm-inference-service] url: str | None = None, [e2e-llm-inference-service] response: BaseHTTPResponse | None = None, [e2e-llm-inference-service] error: Exception | None = None, [e2e-llm-inference-service] _pool: ConnectionPool | None = None, [e2e-llm-inference-service] _stacktrace: TracebackType | None = None, [e2e-llm-inference-service] ) -> Self: [e2e-llm-inference-service] """Return a new Retry object with incremented retry counters. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response: A response object, or None, if the server did not [e2e-llm-inference-service] return a response. [e2e-llm-inference-service] :type response: :class:`~urllib3.response.BaseHTTPResponse` [e2e-llm-inference-service] :param Exception error: An error encountered during the request, or [e2e-llm-inference-service] None if the response was received successfully. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: A new ``Retry`` object. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if self.total is False and error: [e2e-llm-inference-service] # Disabled, indicate to re-raise the error. [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] [e2e-llm-inference-service] total = self.total [e2e-llm-inference-service] if total is not None: [e2e-llm-inference-service] total -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] connect = self.connect [e2e-llm-inference-service] read = self.read [e2e-llm-inference-service] redirect = self.redirect [e2e-llm-inference-service] status_count = self.status [e2e-llm-inference-service] other = self.other [e2e-llm-inference-service] cause = "unknown" [e2e-llm-inference-service] status = None [e2e-llm-inference-service] redirect_location = None [e2e-llm-inference-service] [e2e-llm-inference-service] if error and self._is_connection_error(error): [e2e-llm-inference-service] # Connect retry? [e2e-llm-inference-service] if connect is False: [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif connect is not None: [e2e-llm-inference-service] connect -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error and self._is_read_error(error): [e2e-llm-inference-service] # Read retry? [e2e-llm-inference-service] if read is False or method is None or not self._is_method_retryable(method): [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif read is not None: [e2e-llm-inference-service] read -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error: [e2e-llm-inference-service] # Other retry? [e2e-llm-inference-service] if other is not None: [e2e-llm-inference-service] other -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif response and response.get_redirect_location(): [e2e-llm-inference-service] # Redirect retry? [e2e-llm-inference-service] if redirect is not None: [e2e-llm-inference-service] redirect -= 1 [e2e-llm-inference-service] cause = "too many redirects" [e2e-llm-inference-service] response_redirect_location = response.get_redirect_location() [e2e-llm-inference-service] if response_redirect_location: [e2e-llm-inference-service] redirect_location = response_redirect_location [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Incrementing because of a server error like a 500 in [e2e-llm-inference-service] # status_forcelist and the given method is in the allowed_methods [e2e-llm-inference-service] cause = ResponseError.GENERIC_ERROR [e2e-llm-inference-service] if response and response.status: [e2e-llm-inference-service] if status_count is not None: [e2e-llm-inference-service] status_count -= 1 [e2e-llm-inference-service] cause = ResponseError.SPECIFIC_ERROR.format(status_code=response.status) [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] history = self.history + ( [e2e-llm-inference-service] RequestHistory(method, url, error, status, redirect_location), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] new_retry = self.new( [e2e-llm-inference-service] total=total, [e2e-llm-inference-service] connect=connect, [e2e-llm-inference-service] read=read, [e2e-llm-inference-service] redirect=redirect, [e2e-llm-inference-service] status=status_count, [e2e-llm-inference-service] other=other, [e2e-llm-inference-service] history=history, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if new_retry.is_exhausted(): [e2e-llm-inference-service] reason = error or ResponseError(cause) [e2e-llm-inference-service] > raise MaxRetryError(_pool, url, reason) from reason # type: ignore[arg-type] [e2e-llm-inference-service] E urllib3.exceptions.MaxRetryError: HTTPSConnectionPool(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Max retries exceeded with url: /apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd (Caused by NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Failed to resolve 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/util/retry.py:519: MaxRetryError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig router-managed-rolling-upgrade-64e041bd in namespace kserve-ci-e2e-test [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=2, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Failed to resolve 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)")': /apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=1, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Failed to resolve 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)")': /apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=0, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Failed to resolve 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)")': /apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceserviceconfigs/router-managed-rolling-upgrade-64e041bd [e2e-llm-inference-service] =================================== FAILURES =================================== [e2e-llm-inference-service] _ test_llm_inference_service[router-managed-workload-simulated-dp-ep-cpu-model-fb-opt-125m] _ [e2e-llm-inference-service] [gw0] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-simulated-dp-ep-cpu', 'model-fb-opt-125m'], prompt='This test simulate... {'name': 'model-fb-opt-125m-llmisvc-model-9f2e00e5'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-gateway-ref", [e2e-llm-inference-service] "router-with-managed-route", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[0], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[0]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="custom-route-timeout-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="router-with-refs-test", [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[0], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[0]], [e2e-llm-inference-service] routes=[ROUTER_ROUTES[0], ROUTER_ROUTES[1]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=["router-managed", "workload-pd-cpu", "model-fb-opt-125m"], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="custom-route-timeout-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="router-with-refs-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[1], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[1]], [e2e-llm-inference-service] routes=[ROUTER_ROUTES[2], ROUTER_ROUTES[3]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-dp-ep-gpu", [e2e-llm-inference-service] "workload-dp-ep-prefill-gpu", [e2e-llm-inference-service] "model-deepseek-v2-lite", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="Delve into the multifaceted implications of a fully disaggregated cloud architecture, specifically " [e2e-llm-inference-service] "where the compute plane (P) and the data plane (D) are independently deployed and managed for a " [e2e-llm-inference-service] "geographically distributed, high-throughput, low-latency microservices ecosystem. Beyond the " [e2e-llm-inference-service] "fundamental challenges of network latency and data consistency, elaborate on the advanced " [e2e-llm-inference-service] "considerations and trade-offs inherent in such a setup: 1. Network Architecture and Protocols: " [e2e-llm-inference-service] "How would the network fabric and underlying protocols (e.g., RDMA, custom transport layers) need to " [e2e-llm-inference-service] "evolve to support optimal performance and minimize inter-plane communication overhead, especially for " [e2e-llm-inference-service] "synchronous operations? Discuss the role of network programmability (e.g., SDN, P4) in dynamically " [e2e-llm-inference-service] "optimizing routing and traffic flow between P and D. 2. Advanced Data Consistency and Durability: " [e2e-llm-inference-service] "Explore sophisticated data consistency models (e.g., causal consistency, strong eventual consistency) " [e2e-llm-inference-service] "and their applicability in balancing performance and data integrity across a globally distributed data plane. " [e2e-llm-inference-service] "Detail strategies for ensuring data durability and fault tolerance, including multi-region replication, " [e2e-llm-inference-service] "intelligent partitioning, and recovery mechanisms in the event of partial or full plane failures. " [e2e-llm-inference-service] "3. Dynamic Resource Orchestration and Cost Optimization: Analyze how an orchestration layer would intelligently " [e2e-llm-inference-service] "manage the independent scaling of compute (P) and data (D) resources, considering fluctuating workloads, " [e2e-llm-inference-service] "cost efficiency, and performance targets (e.g., using predictive analytics for resource provisioning). " [e2e-llm-inference-service] "Discuss mechanisms for dynamically reallocating compute nodes to different data partitions based on " [e2e-llm-inference-service] "workload patterns and data locality, potentially involving live migration strategies. " [e2e-llm-inference-service] "4. Security and Compliance in a Distributed Landscape: Address the enhanced security perimeter " [e2e-llm-inference-service] "challenges, including securing communication channels between P and D (encryption in transit, mutual TLS), " [e2e-llm-inference-service] "fine-grained access control to data at rest and in motion, and identity management across disaggregated " [e2e-llm-inference-service] "components. Discuss how such an architecture impacts compliance with regulatory frameworks (e.g., GDPR, HIPAA) " [e2e-llm-inference-service] "concerning data sovereignty, privacy, and auditability. 5. Operational Complexity and Observability: " [e2e-llm-inference-service] "Examine the increased complexity in monitoring, logging, and tracing across highly decoupled compute and " [e2e-llm-inference-service] "data planes. What specialized tooling and practices (e.g., distributed tracing with OpenTelemetry, advanced AIOps) " [e2e-llm-inference-service] "would be essential? How would incident response and troubleshooting differ in this disaggregated environment " [e2e-llm-inference-service] "compared to traditional integrated systems? Consider the challenges of pinpointing root causes across " [e2e-llm-inference-service] "independent failures. 6. Real-world Applicability and Future Trends: Identify specific industries " [e2e-llm-inference-service] "or use cases (e.g., high-frequency trading, IoT edge processing, large language model inference) " [e2e-llm-inference-service] "where the benefits of P/D disaggregation would strongly outweigh its complexities. " [e2e-llm-inference-service] "Conclude by speculating on emerging technologies or paradigms (e.g., serverless compute functions " [e2e-llm-inference-service] "directly interacting with object storage, in-memory disaggregation) that could further drive or " [e2e-llm-inference-service] "transform P/D disaggregation in cloud computing.", [e2e-llm-inference-service] max_tokens=2000, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_gpu, [e2e-llm-inference-service] pytest.mark.cluster_nvidia, [e2e-llm-inference-service] pytest.mark.cluster_nvidia_roce, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-no-scheduler", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.no_scheduler, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="This test simulates DP+EP that can run on CPU, the idea is to test the LWS-based deployment, " [e2e-llm-inference-service] "but without the resources requirements for DP+EP (GPUs and ROCe/IB).", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_multi_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler config tests [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-inline-config-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Chat completions endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] model_name="Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-configmap-ref", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-configmap-ref-test", [e2e-llm-inference-service] before_test=[create_scheduler_configmap], [e2e-llm-inference-service] after_test=[delete_scheduler_configmap], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-replicas", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-ha-replicas-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-custom-template", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-custom-template-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler v0.6 → v0.7 migration tests. [e2e-llm-inference-service] # Deploy v0.6-style configs and verify the controller migrates them [e2e-llm-inference-service] # so the v0.7 scheduler boots successfully. [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-pd-config-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-pd-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-nonzero-threshold-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-threshold-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Precise prefix KV cache routing test [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-precise-prefix-cache-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator-kvcache", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="precise-prefix-cache-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Models endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="data"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/chat/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — LoRA adapter [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] model_name=f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/models (base + LoRA) [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=assert_models_contains( [e2e-llm-inference-service] "facebook/opt-125m", [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] "lora-adapter-1", [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # PVC storage tests -- validate direct PVC volume mount with real vLLM serving [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_multi_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_inference_service(test_case: TestCase): # noqa: F811 [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] if not test_case.llm_service.metadata.annotations: [e2e-llm-inference-service] test_case.llm_service.metadata.annotations = {} [e2e-llm-inference-service] [e2e-llm-inference-service] test_case.llm_service.metadata.annotations[ [e2e-llm-inference-service] "security.opendatahub.io/enable-auth" [e2e-llm-inference-service] ] = "false" [e2e-llm-inference-service] prefix = test_case.log_prefix [e2e-llm-inference-service] [e2e-llm-inference-service] test_failed = False [e2e-llm-inference-service] try: [e2e-llm-inference-service] print(f"{prefix} Creating LLMInferenceService {service_name}") [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] print(f"{prefix} Waiting for LLMInferenceService {service_name} to be ready") [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:812: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...pu-ll-b2c82424'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-llmisvc-model-9f2e00e5'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-02T13:38:32.417489', start_time = 1782999512.4178507 [e2e-llm-inference-service] duration = 900.2282328605652, timestamp_end = '2026-07-02T13:53:32.646105' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....p-ep-cpu-ll-b2c82424'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-llmisvc-model-9f2e00e5'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(condition.get("type")) [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1204: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7fbae4c99580> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1215: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(condition.get("type")) [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] > raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] E AssertionError: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1199: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig router-managed-llmisvc-model-fb-a575f6e7 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig router-managed-llmisvc-model-fb-a575f6e7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig router-managed-llmisvc-model-fb-a575f6e7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig workload-simulated-dp-ep-cpu-ll-b2c82424 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig workload-simulated-dp-ep-cpu-ll-b2c82424 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig workload-simulated-dp-ep-cpu-ll-b2c82424 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig model-fb-opt-125m-llmisvc-model-9f2e00e5 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig model-fb-opt-125m-llmisvc-model-9f2e00e5 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig model-fb-opt-125m-llmisvc-model-9f2e00e5 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_inference_service] [2026-07-02T13:38:32.342486] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-simulated-dp-ep-cpu', 'model-fb-opt-125m'], prompt='This test simulates DP+EP that can run on CPU, the idea is to test the LWS-based deployment, but without the resources requirements for DP+EP (GPUs and ROCe/IB).', service_name='llmisvc-model-fb-opt-125m-route-dc21cb14', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': None, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-fb-opt-125m-route-dc21cb14', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-fb-a575f6e7'}, [e2e-llm-inference-service] {'name': 'workload-simulated-dp-ep-cpu-ll-b2c82424'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-llmisvc-model-9f2e00e5'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-02T13:38:32.356170] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-fb-opt-125m-route-dc21cb14', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-fb-a575f6e7'}, [e2e-llm-inference-service] {'name': 'workload-simulated-dp-ep-cpu-ll-b2c82424'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-llmisvc-model-9f2e00e5'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-02T13:38:32.417305] end - ✅ in 0.061s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-02T13:38:32.417489] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-fb-opt-125m-route-dc21cb14', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-fb-a575f6e7'}, [e2e-llm-inference-service] {'name': 'workload-simulated-dp-ep-cpu-ll-b2c82424'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-llmisvc-model-9f2e00e5'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'Inference Pool kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:38:52Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:38:52Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:38:52Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1219 Timed out waiting: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-02T13:53:32.646105] end - ❌ 900.228s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:831 [router-managed-workload-simulated-dp-ep-cpu-model-fb-opt-125m] ❌ ERROR: Failed to call llm inference service llmisvc-model-fb-opt-125m-route-dc21cb14: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1240 🔍 # Diagnostics for 'llmisvc-model-fb-opt-125m-route-dc21cb14' in 'kserve-ci-e2e-test' [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1241 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1242 # LLMInferenceService llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1245 apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] security.opendatahub.io/enable-auth: 'false' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:32Z' [e2e-llm-inference-service] finalizers: [e2e-llm-inference-service] - serving.kserve.io/llmisvc-finalizer [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:security.opendatahub.io/enable-auth: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:baseRefs: {} [e2e-llm-inference-service] manager: OpenAPI-Generator [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:32Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:finalizers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] v:"serving.kserve.io/llmisvc-finalizer": {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:32Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:addresses: {} [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-router-route: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-scheduler: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-tracing: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-worker-data-parallel: {} [e2e-llm-inference-service] f:appliedConfigs: {} [e2e-llm-inference-service] f:conditions: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:router: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:gateways: {} [e2e-llm-inference-service] f:scheduler: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:inferencePool: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:service: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:url: {} [e2e-llm-inference-service] f:workloads: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:primary: {} [e2e-llm-inference-service] f:scheduler: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:39:12Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] resourceVersion: '55695' [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] baseRefs: [e2e-llm-inference-service] - name: router-managed-llmisvc-model-fb-a575f6e7 [e2e-llm-inference-service] - name: workload-simulated-dp-ep-cpu-ll-b2c82424 [e2e-llm-inference-service] - name: model-fb-opt-125m-llmisvc-model-9f2e00e5 [e2e-llm-inference-service] model: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uri: '' [e2e-llm-inference-service] status: [e2e-llm-inference-service] addresses: [e2e-llm-inference-service] - name: gateway-external-model-routing [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/ [e2e-llm-inference-service] - name: gateway-external [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] - name: gateway-internal-model-routing [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/ [e2e-llm-inference-service] - name: gateway-internal [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-template: kserve-config-llm-decode-template [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-worker-data-parallel: kserve-config-llm-decode-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-template: kserve-config-llm-prefill-template [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-worker-data-parallel: kserve-config-llm-prefill-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-router-route: kserve-config-llm-router-route [e2e-llm-inference-service] serving.kserve.io/config-llm-scheduler: kserve-config-llm-scheduler [e2e-llm-inference-service] serving.kserve.io/config-llm-template: kserve-config-llm-template [e2e-llm-inference-service] serving.kserve.io/config-llm-tracing: kserve-config-llm-tracing [e2e-llm-inference-service] serving.kserve.io/config-llm-worker-data-parallel: kserve-config-llm-worker-data-parallel [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:52Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: HTTPRoutesReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:52Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: InferencePoolReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:39Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: PresetsCombined [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:39:12Z' [e2e-llm-inference-service] message: LWS is progressing [e2e-llm-inference-service] reason: Progressing [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: Ready [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:39:12Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: RouterReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:39:12Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: SchedulerWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:39Z' [e2e-llm-inference-service] message: LWS is progressing [e2e-llm-inference-service] reason: Progressing [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: WorkerWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:39Z' [e2e-llm-inference-service] message: LWS is progressing [e2e-llm-inference-service] reason: Progressing [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: WorkloadsReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:44 TIME NAMESPACE SOURCE TYPE REASON MESSAGE [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:45 -------------------------------------------------------------------------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-6599bc8fd-8sr69 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-6599bc8fd-8sr69 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-router-scheduler-6fcbb55f87 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-6599bc8fd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:52 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy auth-disabled-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "auth-disabled-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-disabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-disabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-disabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-disabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:56 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-disabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-67c57d579b-gtwnj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:05 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 27.649s (27.649s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.35:8000/health": dial tcp 10.132.0.35:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-67c57d579b-gtwnj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.31/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.181s (1.181s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-router-scheduler-688c79755d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-67c57d579b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-enabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-enabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-enabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-enabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-enabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-enabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-6cd6c5884-rrbrr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-6cd6c5884-rrbrr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-router-scheduler-5b958c76d8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-6cd6c5884 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-invalid-token-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-invalid-token-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-invalid-token-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-invalid-token-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-invalid-token-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:24 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-invalid-token-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:25 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.52:8001/health": dial tcp 10.132.0.52:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.53/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.53:8000/health": dial tcp 10.132.0.53:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-router-scheduler-f7968dfdd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-79d6b9cc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:31 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-pd-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-pd-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:43 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-98989556b-vl89r to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:14:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.45:8000/health": dial tcp 10.132.0.45:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-98989556b-vl89r [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-router-scheduler-579b7545d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-98989556b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:57 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:15:01 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/e2e-pvc-model-download-4s522 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:26 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test job-controller Normal SuccessfulCreate Created pod: e2e-pvc-model-download-4s522 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:36 kserve-ci-e2e-test job-controller Normal Completed Job completed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:20 kserve-ci-e2e-test persistentvolume-controller Normal WaitForFirstConsumer waiting for first consumer to be created before binding [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test persistentvolume-controller Normal ExternalProvisioning Waiting for a volume to be created either by the external provisioner 'ebs.csi.aws.com' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal Provisioning External provisioner is provisioning volume for claim "kserve-ci-e2e-test/e2e-pvc-model-storage" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal ProvisioningSucceeded Successfully provisioned volume pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.094s (1.094s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:28 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec0c69dceeb48768325d1a53a749e65786-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.30/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedMount MountVolume.SetUp failed for volume "tls-certs" : secret "gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.60/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:46 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.60:8000/health": dial tcp 10.132.0.60:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.60:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:32 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-3c960099-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-3c960099-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:02 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-3c960099] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" in 808ms (808ms including waiting). Image size: 44914394 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.48:8001/health": dial tcp 10.132.0.48:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.49:8000/health": dial tcp 10.132.0.49:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5b86594bc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:44 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:52 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-50bc673d] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:07:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.44:8000/health": dial tcp 10.132.0.44:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:07:54 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-87882a8e] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 44.193s (44.193s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.62/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:40:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.62:8000/health": dial tcp 10.132.0.62:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 7f6c8cfc6 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:40:17 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test386bd5808c5c450f11fd8632e1f821fb-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:20 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.45:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:27 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-e95b1dc1] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.42:8000/health": dial tcp 10.132.0.42:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler-67f9b884d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:01 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:19 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-7ca60146] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:44 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler-5444b4dbf from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:57 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:55 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-ba4d693a] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.55/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.54/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:25 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 686d468674 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:35 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-2577e794-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:53 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-2577e794] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:56 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.47:8000/health": dial tcp 10.132.0.47:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:51 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-59b9d263-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-59b9d263-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:25 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-59b9d263] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.50:8001/health": dial tcp 10.132.0.50:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.51:8000/health": dial tcp 10.132.0.51:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-79497db4cc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:56 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-e8706282-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:09 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-e8706282] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:30 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.34:8000/health": dial tcp 10.133.0.34:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:36 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:14 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4f8c0978] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.32/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:50 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:48 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-a50492e9] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:40 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-4b931143-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-4b931143-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:21 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-4b931143] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:25 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.42:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-5b1e8f15-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-5b1e8f15-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:22 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-5b1e8f15] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-e45d1f79-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-e45d1f79-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:20 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-e45d1f79] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler-6fcb489785 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.566s (1.566s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-7dbcb75dbc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0, llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc44d181485fad85e662eb092f3749502f-kserve-router-scheduler-69bf65579d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler-5dd88bfbb7 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler-5f555d4d85 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler-54ccf64f4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler-6d86bd4d9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-67b4bb9646 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9, llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler-68cc9685d6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler-749449dbc8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.61/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-single-adapter-test-kserve-85cd6c88dc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-single-adapter-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-single-adapter-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-single-adapter-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-single-adapter-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.29/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.483s (1.483s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.29:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 3.651s (3.651s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.88s (1.88s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 2.962s (2.962s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.723s (1.723s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" in 30.447s (30.447s including waiting). Image size: 2989890188 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Liveness probe failed: timeout: failed to connect service "10.132.0.34:9003" within 1s: context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-router-scheduler-77d88bdcf4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-7c94f9d7c6 from 0 to 2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy precise-prefix-cache-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "precise-prefix-cache-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/precise-prefix-cache-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/precise-prefix-cache-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/precise-prefix-cache-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/precise-prefix-cache-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:15 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [precise-prefix-cache-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:16 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-1-openshift-default-799f46c59b-fnbwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.28/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:31 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "http://10.133.0.28:15021/healthz/ready": context deadline exceeded (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning BackOff Back-off restarting failed container istio-proxy in pod router-gateway-1-openshift-default-799f46c59b-fnbwc_kserve-ci-e2e-test(f244ca92-1cf1-429e-886d-e0d585d0175f) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.28:15021/healthz/ready": dial tcp 10.133.0.28:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-1-openshift-default-799f46c59b-fnbwc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:58 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-1-openshift-default-799f46c59b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:37 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-2-openshift-default-54c789bdc6-w57g9 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:19 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.44:15021/healthz/ready": dial tcp 10.133.0.44:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-2-openshift-default-54c789bdc6-w57g9 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-2-openshift-default-54c789bdc6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:17 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.56:8001/health": dial tcp 10.132.0.56:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.57/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.57:8000/health": dial tcp 10.132.0.57:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-prefill-778968b98b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-router-scheduler-5c684fd54d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-694c6ddf4c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:33 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:47 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-587fcd8566-2pqvl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:20:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:29 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-587fcd8566-2pqvl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-router-scheduler-74545d8489 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-587fcd8566 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:05 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-6hzlt to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.59/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:51 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-f2dmd to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.58/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-f2dmd [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-6hzlt [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:07 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy stop-feature-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "stop-feature-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-stop-feature-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/stop-feature-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/stop-feature-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/stop-feature-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [stop-feature-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.InferencePool kserve-ci-e2e-test/stop-feature-test-inference-pool [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:56 kserve-ci-e2e-test LLMInferenceServiceController Warning LLMInferenceServiceNotReady LLMInferenceService [stop-feature-test] is no longer Ready because of: GatewaysReady, HTTPRoutesReady, InferencePoolReady, MainWorkloadReady, PrefillWorkerWorkloadReady, PrefillWorkloadReady, RouterReady, SchedulerWorkloadReady, WorkerWorkloadReady, WorkloadsReady [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### init-container 'storage-initializer' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 2026-07-02 13:38:37.759 1 storage.initializer INFO [initializer-entrypoint:():17] Initializing, args: (src_uri, dest_path): [('hf://facebook/opt-125m', '/mnt/models')] [e2e-llm-inference-service] 2026-07-02 13:38:37.759 1 storage.initializer INFO [kserve_storage.py:download():166] Copying contents of hf://facebook/opt-125m to local [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/wPaCkH-WbT7GsmxMKKrNZTV4nSM=.ac481c8eb05e4d2496fbe076a38a7b4835dd733d.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_f2f2c4c1-d68c-4148-b6cb-98916e10e708'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/5HHJ6px3_ZRDOG3OxNZMhuycwOk=.a591333512516f58bf2002045dece909a0ccdb8b.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_fc15f0c3-ff70-4d1a-95e2-750b715c82ef'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/Xn7B-BWUGOee2Y6hCZtEhtFu4BE=.38c05904caf6e5b9f04ecda5c973d77e6c1da151.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_ee1a4708-60bf-4a43-b969-3092c0d597a9'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/8_PA_wEVGiVa2goH2H4KQOQpvVY=.b3fb716a3024261980becb2382e31a3780985130.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_edf53ce5-16eb-4853-b4cb-7f0a804cd35b'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/gPcsVCQDYDHk-_n0G9uADl7PXIM=.61c60ec52ed43038fff0fbbd68b080c94b0d94b4c8458dbd65965f9b17631c89.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_3fe0b0a6-2a4d-49f2-9c6f-66f8cfdce564'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/3EVKVggOldJcKSsGjSdoUCN1AyQ=.cf739e3ba86db7791ebab2828cc34b8a5acd3a86.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_3c49f27a-1645-47c1-ac35-99bc1c8f1612'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/PtHk0z_I45atnj23IIRhTExwT3w=.226b0752cac7789c48f0cb3ec53eda48b7be36cc.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_bfbce590-0752-46ea-9ab1-384435cfb144'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/Q1p2l2BzM1m6P5jKvr8WTq1TUio=.2d74da6615135c58cf3cf9ad4cb11e7c613ff9e55fe658a47ab83b6c8d1174a9.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_a99fd9b8-5311-4b12-8bee-df5d5766e2f2'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/ahkChHUJFxEmOdq5GDFEmerRzCY=.5dfa36546b8eddce0e04df3133c30df43fcc3828.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_ad7bba7d-7e35-4d00-9987-29ccfaf1ac31'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/a7eHxRFT3OeMBIFg52k2nfj5m7w=.db7090b0c8b34dd957a7e0656c718f978f9203cc874018f37dda44108be5970a.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_5b9a8470-72cb-4b6b-846f-e4ff99aea42d'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/vzaExXFZNBay89bvlQv-ZcI6BTg=.27c24ca9d908d0b678b20c698aeb9e950c44d865.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_e212ec59-62ee-456b-b18c-4aa8393be3a7'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/j3m-Hy6QvBddw8RXA1uSWl1AJ0c=.0a39732b2d8be8e493cab3da68b68cc3e28221de.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_8cc4c14c-e709-4a2b-a82a-2548afba7038'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] 2026-07-02 13:38:41.539 1 storage.initializer INFO [kserve_storage.py:download():234] Successfully copied hf://facebook/opt-125m to /mnt/models [e2e-llm-inference-service] 2026-07-02 13:38:41.540 1 storage.initializer INFO [kserve_storage.py:download():235] Model downloaded in 3.7802703699999256 seconds. [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:44938 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:44940 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:44956 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:44966 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:44982 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:44998 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:45006 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:45014 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:54916 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:54926 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.132.0.2:54940 - "GET /health HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] (APIServer pid=1) INFO: 10.133.0.49:39420 - "GET /metrics HTTP/1.1" 200 OK [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### init-container 'storage-initializer' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 2026-07-02 13:38:37.882 1 storage.initializer INFO [initializer-entrypoint:():17] Initializing, args: (src_uri, dest_path): [('hf://facebook/opt-125m', '/mnt/models')] [e2e-llm-inference-service] 2026-07-02 13:38:37.882 1 storage.initializer INFO [kserve_storage.py:download():166] Copying contents of hf://facebook/opt-125m to local [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/wPaCkH-WbT7GsmxMKKrNZTV4nSM=.ac481c8eb05e4d2496fbe076a38a7b4835dd733d.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_ce5db927-d9bc-4a61-b21a-7a20a487da26'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/5HHJ6px3_ZRDOG3OxNZMhuycwOk=.a591333512516f58bf2002045dece909a0ccdb8b.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_a8af2681-52c6-4d0d-ba65-bd428058b576'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/Xn7B-BWUGOee2Y6hCZtEhtFu4BE=.38c05904caf6e5b9f04ecda5c973d77e6c1da151.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_18439591-53e3-475e-9d59-8f881f7fe7ef'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/8_PA_wEVGiVa2goH2H4KQOQpvVY=.b3fb716a3024261980becb2382e31a3780985130.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_f10efef3-3873-4f17-83c6-9c52f325d90a'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/gPcsVCQDYDHk-_n0G9uADl7PXIM=.61c60ec52ed43038fff0fbbd68b080c94b0d94b4c8458dbd65965f9b17631c89.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_aced24e3-cf64-45fc-a536-3fdc17b9060d'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/3EVKVggOldJcKSsGjSdoUCN1AyQ=.cf739e3ba86db7791ebab2828cc34b8a5acd3a86.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_7500991d-103b-46cb-ac30-6d57ce3d65b7'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/PtHk0z_I45atnj23IIRhTExwT3w=.226b0752cac7789c48f0cb3ec53eda48b7be36cc.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_dfaf5a2c-6202-4232-a1dc-bf7e858ada08'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/Q1p2l2BzM1m6P5jKvr8WTq1TUio=.2d74da6615135c58cf3cf9ad4cb11e7c613ff9e55fe658a47ab83b6c8d1174a9.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_4858aed1-b1b7-4b49-b3ff-533ebb7e9dec'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/ahkChHUJFxEmOdq5GDFEmerRzCY=.5dfa36546b8eddce0e04df3133c30df43fcc3828.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_04c681cb-d6c3-4293-98be-31124d5dad34'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/a7eHxRFT3OeMBIFg52k2nfj5m7w=.db7090b0c8b34dd957a7e0656c718f978f9203cc874018f37dda44108be5970a.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_050a2e62-79e7-4457-8e57-3171b791c109'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/vzaExXFZNBay89bvlQv-ZcI6BTg=.27c24ca9d908d0b678b20c698aeb9e950c44d865.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_381b2763-73ca-47d2-ab90-3a2db6043586'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/j3m-Hy6QvBddw8RXA1uSWl1AJ0c=.0a39732b2d8be8e493cab3da68b68cc3e28221de.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_af6aa993-2c54-421e-ad71-42b288fbf96a'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] 2026-07-02 13:38:41.655 1 storage.initializer INFO [kserve_storage.py:download():234] Successfully copied hf://facebook/opt-125m to /mnt/models [e2e-llm-inference-service] 2026-07-02 13:38:41.655 1 storage.initializer INFO [kserve_storage.py:download():235] Model downloaded in 3.772961185999975 seconds. [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 INFO 07-02 13:39:39 [importing.py:44] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors. [e2e-llm-inference-service] INFO 07-02 13:39:39 [importing.py:68] Triton not installed or not compatible; certain GPU-related functions will not be available. [e2e-llm-inference-service] (APIServer pid=1) INFO 07-02 13:39:44 [utils.py:299] [e2e-llm-inference-service] (APIServer pid=1) INFO 07-02 13:39:44 [utils.py:299] █ █ █▄ ▄█ [e2e-llm-inference-service] (APIServer pid=1) INFO 07-02 13:39:44 [utils.py:299] ▄▄ ▄█ █ █ █ ▀▄▀ █ version 0.19.0 [e2e-llm-inference-service] (APIServer pid=1) INFO 07-02 13:39:44 [utils.py:299] █▄█▀ █ █ █ █ model /mnt/models [e2e-llm-inference-service] (APIServer pid=1) INFO 07-02 13:39:44 [utils.py:299] ▀▀ ▀▀▀▀▀ ▀▀▀▀▀ ▀ ▀ [e2e-llm-inference-service] (APIServer pid=1) INFO 07-02 13:39:44 [utils.py:299] [e2e-llm-inference-service] (APIServer pid=1) INFO 07-02 13:39:44 [utils.py:233] non-default args: {'model_tag': '/mnt/models', 'ssl_keyfile': '/var/run/kserve/tls/tls.key', 'ssl_certfile': '/var/run/kserve/tls/tls.crt', 'enable_ssl_refresh': True, 'model': '/mnt/models', 'served_model_name': ['facebook/opt-125m']} [e2e-llm-inference-service] (APIServer pid=1) WARNING 07-02 13:39:44 [arg_utils.py:1390] The global random seed is set to 0. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM. [e2e-llm-inference-service] (APIServer pid=1) INFO 07-02 13:40:02 [model.py:549] Resolved architecture: OPTForCausalLM [e2e-llm-inference-service] (APIServer pid=1) INFO 07-02 13:40:02 [model.py:1678] Using max model len 2048 [e2e-llm-inference-service] (APIServer pid=1) INFO 07-02 13:40:02 [vllm.py:790] Asynchronous scheduling is enabled. [e2e-llm-inference-service] INFO 07-02 13:40:15 [importing.py:44] Triton is installed but 0 active driver(s) found (expected 1). Disabling Triton to prevent runtime errors. [e2e-llm-inference-service] INFO 07-02 13:40:15 [importing.py:68] Triton not installed or not compatible; certain GPU-related functions will not be available. [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [core.py:105] Initializing a V1 LLM engine (v0.19.0) with config: model='/mnt/models', speculative_config=None, tokenizer='/mnt/models', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=False, dtype=torch.float16, max_seq_len=2048, download_dir=None, load_format=auto, tensor_parallel_size=1, pipeline_parallel_size=1, data_parallel_size=1, decode_context_parallel_size=1, dcp_comm_backend=ag_rs, disable_custom_all_reduce=True, quantization=None, enforce_eager=False, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cpu, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=0, served_model_name=facebook/opt-125m, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'mode': , 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['none'], 'splitting_ops': [], 'compile_mm_encoder': False, 'cudagraph_mm_encoder': False, 'encoder_cudagraph_token_budgets': [], 'encoder_cudagraph_max_images_per_batch': 0, 'compile_sizes': None, 'compile_ranges_endpoints': [2048], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'size_asserts': False, 'alignment_asserts': False, 'scalar_asserts': False, 'dce': True, 'nan_asserts': False, 'epilogue_fusion': True, 'cpp.dynamic_threads': True}, 'inductor_passes': {}, 'cudagraph_mode': , 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False}, 'max_cudagraph_capture_size': None, 'dynamic_shapes_config': {'type': , 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [cpu_worker.py:236] auto thread-binding list (id, physical core): [(4, 0), (5, 1), (6, 2), (7, 3)] [e2e-llm-inference-service] [W702 13:40:20.406699472 utils.cpp:76] Warning: numa_migrate_pages failed. errno: 1 (function init_cpu_threads_env) [e2e-llm-inference-service] [W702 13:40:20.406723462 utils.cpp:103] Warning: NUMA binding: Using MEMBIND policy for memory allocation on the NUMA nodes (0). Memory allocations will be strictly bound to these NUMA nodes. (function init_cpu_threads_env) [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [cpu_worker.py:109] OMP threads binding of Process 41: [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [cpu_worker.py:109] OMP tid: 41, core 4 [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [cpu_worker.py:109] OMP tid: 58, core 5 [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [cpu_worker.py:109] OMP tid: 59, core 6 [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [cpu_worker.py:109] OMP tid: 60, core 7 [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [cpu_worker.py:109] [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [parallel_state.py:1400] world_size=1 rank=0 local_rank=0 distributed_init_method=tcp://10.133.0.48:48187 backend=gloo [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] [Gloo] Rank 0 is connected to 0 peer ranks. Expected number of connected peer ranks is : 0 [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [parallel_state.py:1716] rank 0 in world size 1 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A, EPLB rank N/A [e2e-llm-inference-service] (EngineCore pid=41) INFO 07-02 13:40:20 [cpu_model_runner.py:71] Starting to load model /mnt/models... [e2e-llm-inference-service] (EngineCore pid=41) Loading pt checkpoint shards: 0% Completed | 0/1 [00:00():17] Initializing, args: (src_uri, dest_path): [('hf://facebook/opt-125m', '/mnt/models')] [e2e-llm-inference-service] 2026-07-02 13:38:37.984 1 storage.initializer INFO [kserve_storage.py:download():166] Copying contents of hf://facebook/opt-125m to local [e2e-llm-inference-service] 2026-07-02 13:38:37.984 1 storage.initializer INFO [kserve_storage.py:download():169] Allow patterns: ['tokenizer.json', 'tokenizer_config.json', 'special_tokens_map.json', 'vocab.json', 'merges.txt', 'config.json', 'generation_config.json'] [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/8_PA_wEVGiVa2goH2H4KQOQpvVY=.b3fb716a3024261980becb2382e31a3780985130.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_a64fad78-e34e-4db1-820f-31159becb32f'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/3EVKVggOldJcKSsGjSdoUCN1AyQ=.cf739e3ba86db7791ebab2828cc34b8a5acd3a86.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_c74940fd-9755-4a27-b711-50cdf1df56ba'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/PtHk0z_I45atnj23IIRhTExwT3w=.226b0752cac7789c48f0cb3ec53eda48b7be36cc.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_20dd0513-893a-4909-a540-570666fe44e3'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/ahkChHUJFxEmOdq5GDFEmerRzCY=.5dfa36546b8eddce0e04df3133c30df43fcc3828.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_ee3ef742-a209-47cd-8b42-d6ca0ff18af4'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/vzaExXFZNBay89bvlQv-ZcI6BTg=.27c24ca9d908d0b678b20c698aeb9e950c44d865.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_d9a82bd3-24fb-4db0-a6a1-f5f2628bf993'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/j3m-Hy6QvBddw8RXA1uSWl1AJ0c=.0a39732b2d8be8e493cab3da68b68cc3e28221de.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_3527678c-b9a9-4237-97e4-b6a0c16dc239'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] 2026-07-02 13:38:38.424 1 storage.initializer INFO [kserve_storage.py:download():234] Successfully copied hf://facebook/opt-125m to /mnt/models [e2e-llm-inference-service] 2026-07-02 13:38:38.424 1 storage.initializer INFO [kserve_storage.py:download():235] Model downloaded in 0.44007728199994745 seconds. [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 {"level":"info","ts":1782999518.694432,"logger":"setup","caller":"runner/runner.go:196","msg":"GIE build","commit-sha":"181aa8358916e19b8844ccc752b2d6153d4b2ad6","build-ref":"v0.9.0-rc.2"} [e2e-llm-inference-service] Flag --model-server-metrics-scheme has been deprecated, This flag is deprecated. Configure via EndpointPickerConfig data layer plugin parameters instead. [e2e-llm-inference-service] {"level":"info","ts":1782999518.6945858,"logger":"setup","caller":"runner/runner.go:217","msg":"Flags processed","flags":{"cert-path":"/var/run/kserve/tls","config-file":"","config-text":"apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type: prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n- name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n","disable-endpoint-subset-filter":false,"enable-cert-reload":true,"enable-grpc-stream-metrics":false,"enable-pprof":true,"endpoint-selector":"","endpoint-target-ports":{},"grpc-health-port":9003,"grpc-max-recv-msg-size":"","grpc-max-send-msg-size":"","grpc-port":9002,"ha-enable-leader-election":false,"health-checking":false,"metrics-endpoint-auth":true,"metrics-port":9090,"metrics-staleness-threshold":2000000000,"model-server-metrics-https-insecure-skip-verify":true,"model-server-metrics-path":"/metrics","model-server-metrics-port":0,"model-server-metrics-scheme":"https","pool-group":"inference.networking.k8s.io","pool-name":"llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool","pool-namespace":"kserve-ci-e2e-test","refresh-metrics-interval":50000000,"refresh-prometheus-metrics-interval":5000000000,"secure-serving":true,"tracing":true,"v":2,"zap-devel":{},"zap-encoder":{},"zap-log-level":{},"zap-stacktrace-level":{},"zap-time-encoding":{}}} [e2e-llm-inference-service] {"level":"info","ts":1782999518.6947403,"logger":"setup.trace","caller":"tracing/telemetry.go:123","msg":"init OTel trace exporter","type":"console"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.6954195,"caller":"loader/configloader.go:89","msg":"DEPRECATION: apiVersion inference.networking.x-k8s.io/v1alpha1/EndpointPickerConfig is deprecated","replacement":"llm-d.ai/v1alpha1/EndpointPickerConfig"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.6954606,"caller":"loader/configloader.go:121","msg":"Loaded raw configuration","config":"{Plugins: [{Type: single-profile-handler} {Type: queue-scorer} {Type: prefix-cache-scorer} {Type: max-score-picker}], SchedulingProfiles: [{Name: default, Plugins: [{PluginRef: queue-scorer, Weight: 2.00} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: max-score-picker}]}]}"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.6954703,"logger":"setup","caller":"runner/runner.go:622","msg":"Data layer: ENABLED"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.6958206,"logger":"setup","caller":"runner/runner.go:281","msg":"Raw config after phase one","config":{"apiVersion":"inference.networking.x-k8s.io/v1alpha1","dataLayer":null,"kind":"EndpointPickerConfig","plugins":[{"name":"single-profile-handler","parameters":null,"type":"single-profile-handler"},{"name":"queue-scorer","parameters":null,"type":"queue-scorer"},{"name":"prefix-cache-scorer","parameters":null,"type":"prefix-cache-scorer"},{"name":"max-score-picker","parameters":null,"type":"max-score-picker"}],"schedulingProfiles":[{"name":"default","plugins":[{"pluginRef":"queue-scorer","weight":2},{"pluginRef":"prefix-cache-scorer","weight":3},{"pluginRef":"max-score-picker","weight":null}]}]}} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7239027,"logger":"utilization-detector/utilization-detector","caller":"utilization/detector.go:83","msg":"Creating new UtilizationDetector","queueDepthThreshold":5,"kvCacheUtilThreshold":0.8,"metricsStalenessThreshold":"200ms","headroom":0} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7240107,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"vllm","mapping":"Mapping{all specs enabled}"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7240632,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"sglang","mapping":"Mapping{disabled: [lora]}"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7241013,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"trtllm-serve","mapping":"Mapping{disabled: [lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.724198,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"triton-tensorrt-llm","mapping":"Mapping{disabled: [lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.724222,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"triton","mapping":"Mapping{disabled: [kv, lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7242908,"caller":"loader/configloader.go:154","msg":"Instantiated all plugins and applied system defaults. Effective raw configuration","config":"{Plugins: [{Name: single-profile-handler, Type: single-profile-handler} {Name: queue-scorer, Type: queue-scorer} {Name: prefix-cache-scorer, Type: prefix-cache-scorer} {Name: max-score-picker, Type: max-score-picker} {Name: fcfs-ordering-policy, Type: fcfs-ordering-policy} {Name: global-strict-fairness-policy, Type: global-strict-fairness-policy} {Name: static-usage-limit-policy, Type: static-usage-limit-policy} {Name: openai-parser, Type: openai-parser} {Name: anthropic-parser, Type: anthropic-parser} {Name: vllmhttp-parser, Type: vllmhttp-parser} {Name: utilization-detector, Type: utilization-detector} {Name: metrics-data-source, Type: metrics-data-source} {Name: core-metrics-extractor, Type: core-metrics-extractor}], SchedulingProfiles: [{Name: default, Plugins: [{PluginRef: queue-scorer, Weight: 2.00} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: max-score-picker}]}], DataLayer: {Sources: [{PluginRef: metrics-data-source, Extractors: [{PluginRef: core-metrics-extractor}]}], Discovery: }, FlowControl: {MaxBytes: unlimited, MaxRequests: unlimited, SaturationDetector: {PluginRef: utilization-detector}}, RequestHandler: {Parsers: [{PluginRef: openai-parser}, {PluginRef: anthropic-parser}, {PluginRef: vllmhttp-parser}]}}"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7243667,"caller":"approximateprefix/plugin.go:88","msg":"Prefix DataProducer initialized","config":{"autoTune":true,"blockSizeTokens":16,"blockSize":0,"maxPrefixBlocksToMatch":2048,"maxPrefixTokensToMatch":131072,"lruCapacityPerServer":31250}} [e2e-llm-inference-service] {"level":"info","ts":1782999518.724459,"caller":"approximateprefix/plugin.go:111","msg":"WARNING: configured blockSizeTokens is below the recommended minimum, overriding it.","blockSizeTokens":16,"minimum":64,"issue":"https://github.com/llm-d/llm-d-router/issues/1158"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7244918,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"approx-prefix-cache-producer/approx-prefix-cache-producer","dataKey":"PrefixCacheMatchInfoDataKey/approx-prefix-cache-producer","consumer":"prefix-cache-scorer"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7245219,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"token-producer/token-producer","dataKey":"TokenizedPrompt/token-producer","consumer":"approx-prefix-cache-producer"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7246063,"caller":"runner/runner.go:685","msg":"loaded configuration from file/text successfully"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7246213,"logger":"setup","caller":"runner/runner.go:308","msg":"EPP config after phase two","config":"{SchedulerConfig:{ProfileHandler: single-profile-handler/single-profile-handler, Profiles: map[default:{Filters: [], Scorers: [queue-scorer/queue-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000], Picker: max-score-picker/max-score-picker}]} SaturationDetector:0xc0008ac4c0 DataConfig:{Sources:[{Plugin:0xc0008b8090 Extractors:[0xc0008ac6c0]}]} FlowControlConfig: ParserRegistry:0xc0008acb40}"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7481868,"logger":"setup","caller":"runner/runner.go:352","msg":"Setting pprof handlers"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7482245,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/block"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7482452,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/cmdline"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7482507,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/profile"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7482553,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/symbol"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7482595,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/trace"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7482646,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/goroutine"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7482688,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/threadcreate"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7482734,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/mutex"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.748279,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7482836,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/heap"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7482882,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/allocs"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7483032,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/plugins/state"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7483106,"logger":"setup","caller":"runner/runner.go:373","msg":"parsed config","scheduler-config":"{ProfileHandler: single-profile-handler/single-profile-handler, Profiles: map[default:{Filters: [], Scorers: [queue-scorer/queue-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000], Picker: max-score-picker/max-score-picker}]}"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7483354,"logger":"setup","caller":"datalayer/runtime.go:99","msg":"Configuring datalayer runtime","numSources":1} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7483444,"logger":"setup","caller":"datalayer/runtime.go:118","msg":"Processing source","source":"metrics-data-source","numExtractors":1} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7483587,"logger":"setup","caller":"datalayer/runtime.go:147","msg":"Source configured","source":"metrics-data-source","extractors":["core-metrics-extractor/core-metrics-extractor"]} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7483723,"logger":"setup","caller":"datalayer/runtime.go:206","msg":"Datalayer runtime configured","pollers":1,"notifiers":0,"endpointSources":0} [e2e-llm-inference-service] {"level":"info","ts":1782999518.748381,"logger":"setup","caller":"runner/runner.go:833","msg":"Experimental Flow Control layer is disabled, using legacy admission control"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7484577,"logger":"setup","caller":"runner/runner.go:721","msg":"ExtProc server runner added to manager."} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7484682,"logger":"setup","caller":"runner/runner.go:260","msg":"Controller manager starting"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7485085,"logger":"controller-runtime.metrics","caller":"server/server.go:208","msg":"Starting metrics server"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.748905,"logger":"controller-runtime.metrics","caller":"server/server.go:247","msg":"Serving metrics server","bindAddress":":9090","secure":false} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7491,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"health"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7492216,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"health","port":9003} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7495818,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","source":"kind source: *v1.InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7499769,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite","source":"kind source: *v1alpha2.InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.750186,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective","source":"kind source: *v1alpha2.InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7506778,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"pod","controllerGroup":"","controllerKind":"Pod","source":"kind source: *v1.Pod"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7508948,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"ext-proc"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7511587,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"ext-proc","port":9002} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7574215,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceObjective","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7574215,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceModelRewrite","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.759319,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.Pod","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.7609675,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.InferencePool","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.8510106,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.851046,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1782999518.8521025,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.8521237,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1782999518.9515853,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.9516213,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1782999518.9517612,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool","reconcileID":"c2c65f38-c5e2-47d1-ad5a-8f20c942cf9f"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.9526207,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"pod","controllerGroup":"","controllerKind":"Pod"} [e2e-llm-inference-service] {"level":"info","ts":1782999518.9526446,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"pod","controllerGroup":"","controllerKind":"Pod","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1782999531.409984,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool","reconcileID":"d1468eaa-889c-4189-9ebe-fe023818bfe2"} [e2e-llm-inference-service] {"level":"info","ts":1782999617.3409524,"caller":"controller/pod_reconciler.go:99","msg":"Pod already exists","controller":"pod","controllerGroup":"","controllerKind":"Pod","Pod":{"name":"llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0","reconcileID":"84a4629a-c43e-4c48-a173-1f2d27eac9fa"} [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 25c7b42c-86ab-40b4-b785-433a4bcabed2 [e2e-llm-inference-service] resourceVersion: '55682' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.133.0.49 [e2e-llm-inference-service] nodeName: ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] uid: 24584ac8-f2d7-4ff1-a230-2cca9b207c69 [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 423893cf-cc3b-40e5-8154-ed30beac9f84 [e2e-llm-inference-service] resourceVersion: '56611' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T13:40:17Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:40:17Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.132.0.62 [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] uid: b57ab898-fcea-428c-8565-f6e86b33169a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] generateName: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: b57ab898-fcea-428c-8565-f6e86b33169a [e2e-llm-inference-service] resourceVersion: '56610' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-leader [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] apps.kubernetes.io/pod-index: '0' [e2e-llm-inference-service] controller-revision-hash: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-57dd9f69d [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: 269370d6a88f087dbcb80280bf4153dc63eb8499 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7f6c8cfc6 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '0' [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] statefulset.kubernetes.io/pod-name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.132.0.62/23"],"mac_address":"0a:58:0a:84:00:3e","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.62/23","gateway_ip":"10.132.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.132.0.62\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:84:00:3e\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/size: '2' [e2e-llm-inference-service] openshift.io/scc: openshift-ai-llminferenceservice-scc [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: serviceaccount [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: StatefulSet [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] uid: f3336882-f5dc-4b47-913d-3bfc0ae35166 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-134-1 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/size: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:apps.kubernetes.io/pod-index: {} [e2e-llm-inference-service] f:controller-revision-hash: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/name: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/template-revision-hash: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/worker-index: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:statefulset.kubernetes.io/pod-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"f3336882-f5dc-4b47-913d-3bfc0ae35166"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:add: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:hostname: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:subdomain: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:40:17Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:initContainerStatuses: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.132.0.62"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kube-api-access-sjszs [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: LWS_LEADER_ADDRESS [e2e-llm-inference-service] value: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0.llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn.kserve-ci-e2e-test [e2e-llm-inference-service] - name: LWS_GROUP_SIZE [e2e-llm-inference-service] value: '2' [e2e-llm-inference-service] - name: LWS_WORKER_INDEX [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-sjszs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: LWS_LEADER_ADDRESS [e2e-llm-inference-service] value: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0.llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn.kserve-ci-e2e-test [e2e-llm-inference-service] - name: LWS_GROUP_SIZE [e2e-llm-inference-service] value: '2' [e2e-llm-inference-service] - name: LWS_WORKER_INDEX [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-sjszs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] serviceAccount: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: llmisvc-model-fb-opt-125m-route-dc21cb-57cf7736-dockercfg-xmmxt [e2e-llm-inference-service] hostname: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] subdomain: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:38:41Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:40:17Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:40:17Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] hostIP: 10.0.134.1 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.134.1 [e2e-llm-inference-service] podIP: 10.132.0.62 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.132.0.62 [e2e-llm-inference-service] startTime: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] initContainerStatuses: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] state: [e2e-llm-inference-service] terminated: [e2e-llm-inference-service] exitCode: 0 [e2e-llm-inference-service] reason: Completed [e2e-llm-inference-service] startedAt: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] finishedAt: '2026-07-02T13:38:41Z' [e2e-llm-inference-service] containerID: cri-o://178f3aee61a0ecbbf9455d63e88f59e2c15c02ae0ccc25e59429dd6daab53538 [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] imageID: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] containerID: cri-o://178f3aee61a0ecbbf9455d63e88f59e2c15c02ae0ccc25e59429dd6daab53538 [e2e-llm-inference-service] started: false [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-sjszs [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T13:38:42Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] imageID: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0 [e2e-llm-inference-service] containerID: cri-o://7393781126e4689057d7f93a4f323322b5c148646847a1d8d983a17d5b3a9200 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-sjszs [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 [e2e-llm-inference-service] generateName: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 00db9fbf-206e-47c5-b6c0-a0722bb83a7b [e2e-llm-inference-service] resourceVersion: '55920' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] apps.kubernetes.io/pod-index: '1' [e2e-llm-inference-service] controller-revision-hash: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-5c59c86cbf [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: 269370d6a88f087dbcb80280bf4153dc63eb8499 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7f6c8cfc6 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '1' [e2e-llm-inference-service] statefulset.kubernetes.io/pod-name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.133.0.48/23"],"mac_address":"0a:58:0a:85:00:30","gateway_ips":["10.133.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.133.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.133.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.133.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.133.0.1"}],"ip_address":"10.133.0.48/23","gateway_ip":"10.133.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.133.0.48\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:85:00:30\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/leader-name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/size: '2' [e2e-llm-inference-service] openshift.io/scc: openshift-ai-llminferenceservice-scc [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: serviceaccount [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: StatefulSet [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] uid: a1ef0a80-adfa-4bff-8922-83957e5cb85e [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/leader-name: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/size: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:apps.kubernetes.io/pod-index: {} [e2e-llm-inference-service] f:controller-revision-hash: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/group-index: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/group-key: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/name: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/template-revision-hash: {} [e2e-llm-inference-service] f:statefulset.kubernetes.io/pod-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"a1ef0a80-adfa-4bff-8922-83957e5cb85e"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:add: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:hostname: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:subdomain: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: ip-10-0-141-170 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:39:27Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:initContainerStatuses: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.133.0.48"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kube-api-access-hq6gs [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: LWS_LEADER_ADDRESS [e2e-llm-inference-service] value: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0.llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn.kserve-ci-e2e-test [e2e-llm-inference-service] - name: LWS_GROUP_SIZE [e2e-llm-inference-service] value: '2' [e2e-llm-inference-service] - name: LWS_WORKER_INDEX [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-hq6gs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: LWS_LEADER_ADDRESS [e2e-llm-inference-service] value: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0.llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn.kserve-ci-e2e-test [e2e-llm-inference-service] - name: LWS_GROUP_SIZE [e2e-llm-inference-service] value: '2' [e2e-llm-inference-service] - name: LWS_WORKER_INDEX [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-hq6gs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] serviceAccount: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] nodeName: ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: llmisvc-model-fb-opt-125m-route-dc21cb-57cf7736-dockercfg-xmmxt [e2e-llm-inference-service] hostname: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 [e2e-llm-inference-service] subdomain: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:38:38Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:38:42Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:39:27Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:39:27Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] hostIP: 10.0.141.170 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.141.170 [e2e-llm-inference-service] podIP: 10.133.0.48 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.133.0.48 [e2e-llm-inference-service] startTime: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] initContainerStatuses: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] state: [e2e-llm-inference-service] terminated: [e2e-llm-inference-service] exitCode: 0 [e2e-llm-inference-service] reason: Completed [e2e-llm-inference-service] startedAt: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] finishedAt: '2026-07-02T13:38:41Z' [e2e-llm-inference-service] containerID: cri-o://18ffe6d33248b20a72a5d41fc4953948a42f874e719f939bfd542258f751ba33 [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] imageID: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] containerID: cri-o://18ffe6d33248b20a72a5d41fc4953948a42f874e719f939bfd542258f751ba33 [e2e-llm-inference-service] started: false [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-hq6gs [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T13:39:26Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] imageID: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0 [e2e-llm-inference-service] containerID: cri-o://d72a6b60481cc763d259d1bdb6638d79d5bf1b80daba0df428bc91e849f03daf [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-hq6gs [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] generateName: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 24584ac8-f2d7-4ff1-a230-2cca9b207c69 [e2e-llm-inference-service] resourceVersion: '55680' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 5688c7c666 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.133.0.49/23"],"mac_address":"0a:58:0a:85:00:31","gateway_ips":["10.133.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.133.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.133.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.133.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.133.0.1"}],"ip_address":"10.133.0.49/23","gateway_ip":"10.133.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.133.0.49\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:85:00:31\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666 [e2e-llm-inference-service] uid: d83e71f7-8e24-429c-b7a5-86012dc74421 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-141-170 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d83e71f7-8e24-429c-b7a5-86012dc74421"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"STORAGE_ALLOW_PATTERNS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:initContainerStatuses: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.133.0.49"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kube-api-access-z27vm [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] - name: STORAGE_ALLOW_PATTERNS [e2e-llm-inference-service] value: '["tokenizer.json", "tokenizer_config.json", "special_tokens_map.json", [e2e-llm-inference-service] "vocab.json", "merges.txt", "config.json", "generation_config.json"]' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-z27vm [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type: prefix-cache-scorer\n\ [e2e-llm-inference-service] - type: max-score-picker\nschedulingProfiles:\n- name: default\n plugins:\n\ [e2e-llm-inference-service] \ - pluginRef: queue-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n\ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-z27vm [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] nodeName: ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] fsGroup: 1000690000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa-dockercfg-l5xfh [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:38:38Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:38:38Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] hostIP: 10.0.141.170 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.141.170 [e2e-llm-inference-service] podIP: 10.133.0.49 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.133.0.49 [e2e-llm-inference-service] startTime: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] initContainerStatuses: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] state: [e2e-llm-inference-service] terminated: [e2e-llm-inference-service] exitCode: 0 [e2e-llm-inference-service] reason: Completed [e2e-llm-inference-service] startedAt: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] finishedAt: '2026-07-02T13:38:38Z' [e2e-llm-inference-service] containerID: cri-o://84a0616e31b36e7c9ec21a1c9ce68d3080b407ca731b560a8c01cd073919384f [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] imageID: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] containerID: cri-o://84a0616e31b36e7c9ec21a1c9ce68d3080b407ca731b560a8c01cd073919384f [e2e-llm-inference-service] started: false [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-z27vm [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T13:38:38Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] imageID: ghcr.io/llm-d/llm-d-router-endpoint-picker@sha256:06b6c75d77afd0e07053402752a9736c2dfbc12a306d0d37d963aac4c1d4e6a6 [e2e-llm-inference-service] containerID: cri-o://773a1d7dc19e13563e9090b827a5d993e41aaa378498b93ba783d6dd71412db7 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-z27vm [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: e6d3e57d-6521-4d9a-a15c-c4f861bdea25 [e2e-llm-inference-service] resourceVersion: '55030' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa-dockercfg-l5xfh [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa-dockercfg-l5xfh"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-x9l7t"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa-dockercfg-l5xfh [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa-dockercfg-l5xfh [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 45898698-1106-4eb0-86e1-52a8b07f07fc [e2e-llm-inference-service] resourceVersion: '55008' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: llmisvc-model-fb-opt-125m-route-dc21cb-57cf7736-dockercfg-xmmxt [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"llmisvc-model-fb-opt-125m-route-dc21cb-57cf7736-dockercfg-xmmxt"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-x9l7t"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: llmisvc-model-fb-opt-125m-route-dc21cb-57cf7736-dockercfg-xmmxt [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: llmisvc-model-fb-opt-125m-route-dc21cb-57cf7736-dockercfg-xmmxt [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 2573e5cb-08c2-4518-9ba1-4fab810a6a81 [e2e-llm-inference-service] resourceVersion: '55068' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] targetPort: grpc [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] targetPort: grpc-health [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] targetPort: metrics [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] targetPort: zmq [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] clusterIP: 172.31.95.129 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.95.129 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: bceb5266-96b2-4de9-8b6c-d7504e1a0b08 [e2e-llm-inference-service] resourceVersion: '55020' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] targetPort: 8000 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] clusterIP: 172.31.73.65 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.73.65 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-5c59c86cbf [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 5d145875-63a0-467f-87e3-f5f086139a66 [e2e-llm-inference-service] resourceVersion: '55047' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] controller.kubernetes.io/hash: 5c59c86cbf [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: 269370d6a88f087dbcb80280bf4153dc63eb8499 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7f6c8cfc6 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: StatefulSet [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] uid: a1ef0a80-adfa-4bff-8922-83957e5cb85e [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:data: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:controller.kubernetes.io/hash: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/group-index: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/group-key: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/name: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/template-revision-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"a1ef0a80-adfa-4bff-8922-83957e5cb85e"}: {} [e2e-llm-inference-service] f:revision: {} [e2e-llm-inference-service] data: [e2e-llm-inference-service] spec: [e2e-llm-inference-service] template: [e2e-llm-inference-service] $patch: replace [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/leader-name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/size: '2' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: 269370d6a88f087dbcb80280bf4153dc63eb8499 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7f6c8cfc6 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] name: main [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /home [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - mountPath: /tmp [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - mountPath: /dev/shm [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - mountPath: /models [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - mountPath: /var/run/kserve/tls [e2e-llm-inference-service] name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] name: storage-initializer [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] serviceAccount: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] serviceAccountName: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] secretName: llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] revision: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ControllerRevision [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-57dd9f69d [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: fd7fdf2e-9a44-4fe8-a657-18d347ba85e5 [e2e-llm-inference-service] resourceVersion: '55027' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-leader [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] controller.kubernetes.io/hash: 57dd9f69d [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7f6c8cfc6 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '0' [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/replicas: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: StatefulSet [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] uid: f3336882-f5dc-4b47-913d-3bfc0ae35166 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:data: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/replicas: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:controller.kubernetes.io/hash: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/name: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/template-revision-hash: {} [e2e-llm-inference-service] f:leaderworkerset.sigs.k8s.io/worker-index: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"f3336882-f5dc-4b47-913d-3bfc0ae35166"}: {} [e2e-llm-inference-service] f:revision: {} [e2e-llm-inference-service] data: [e2e-llm-inference-service] spec: [e2e-llm-inference-service] template: [e2e-llm-inference-service] $patch: replace [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/size: '2' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-leader [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7f6c8cfc6 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '0' [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] spec: [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] name: main [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /home [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - mountPath: /tmp [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - mountPath: /dev/shm [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - mountPath: /models [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - mountPath: /var/run/kserve/tls [e2e-llm-inference-service] name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] name: storage-initializer [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] serviceAccount: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] serviceAccountName: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] secretName: llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] revision: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ControllerRevision [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: c877159a-d382-4129-84a7-4b2846394540 [e2e-llm-inference-service] resourceVersion: '55684' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"STORAGE_ALLOW_PATTERNS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] - name: STORAGE_ALLOW_PATTERNS [e2e-llm-inference-service] value: '["tokenizer.json", "tokenizer_config.json", "special_tokens_map.json", [e2e-llm-inference-service] "vocab.json", "merges.txt", "config.json", "generation_config.json"]' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n-\ [e2e-llm-inference-service] \ name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: Recreate [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666" [e2e-llm-inference-service] has successfully progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: d83e71f7-8e24-429c-b7a5-86012dc74421 [e2e-llm-inference-service] resourceVersion: '55683' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 5688c7c666 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler [e2e-llm-inference-service] uid: c877159a-d382-4129-84a7-4b2846394540 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"c877159a-d382-4129-84a7-4b2846394540"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"STORAGE_ALLOW_PATTERNS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 5688c7c666 [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 5688c7c666 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] - name: STORAGE_ALLOW_PATTERNS [e2e-llm-inference-service] value: '["tokenizer.json", "tokenizer_config.json", "special_tokens_map.json", [e2e-llm-inference-service] "vocab.json", "merges.txt", "config.json", "generation_config.json"]' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n-\ [e2e-llm-inference-service] \ name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-rb [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 56b12e95-d2de-48d1-859a-3ab3f48e2c5e [e2e-llm-inference-service] resourceVersion: '55061' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-scc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: df961ffb-2448-4fb8-82b2-f6763ed56006 [e2e-llm-inference-service] resourceVersion: '55005' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-prefill [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: ClusterRole [e2e-llm-inference-service] name: openshift-ai-llminferenceservice-scc [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 70870f2f-9932-4284-8adb-9e6b1df115d7 [e2e-llm-inference-service] resourceVersion: '55041' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] - create [e2e-llm-inference-service] - update [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - delete [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service-mfbmd [e2e-llm-inference-service] generateName: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: c4753d74-8491-4ec0-bb0f-deecf9ab9c38 [e2e-llm-inference-service] resourceVersion: '55681' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service [e2e-llm-inference-service] uid: 2573e5cb-08c2-4518-9ba1-4fab810a6a81 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:39:11Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"2573e5cb-08c2-4518-9ba1-4fab810a6a81"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.133.0.49 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] uid: 24584ac8-f2d7-4ff1-a230-2cca9b207c69 [e2e-llm-inference-service] nodeName: ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-sz9mgh [e2e-llm-inference-service] generateName: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 3ca46ac6-911d-4c58-aae8-3b108e3787e6 [e2e-llm-inference-service] resourceVersion: '56614' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T13:40:17Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] uid: bceb5266-96b2-4de9-8b6c-d7504e1a0b08 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:40:17Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"bceb5266-96b2-4de9-8b6c-d7504e1a0b08"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.132.0.62 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] uid: b57ab898-fcea-428c-8565-f6e86b33169a [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-rb [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 56b12e95-d2de-48d1-859a-3ab3f48e2c5e [e2e-llm-inference-service] resourceVersion: '55061' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:kserve-ci-e2e-test:llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-scc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: df961ffb-2448-4fb8-82b2-f6763ed56006 [e2e-llm-inference-service] resourceVersion: '55005' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:kserve-ci-e2e-test:llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] - system:serviceaccount:kserve-ci-e2e-test:llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-prefill [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-prefill [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] name: openshift-ai-llminferenceservice-scc [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 70870f2f-9932-4284-8adb-9e6b1df115d7 [e2e-llm-inference-service] resourceVersion: '55041' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - create [e2e-llm-inference-service] - delete [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - update [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:51Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:38:51Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:38:52Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55400' [e2e-llm-inference-service] uid: 15b58bbf-994a-41a6-9151-cab35f73a097 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:51Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:51Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:38:51Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:38:52Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55400' [e2e-llm-inference-service] uid: 15b58bbf-994a-41a6-9151-cab35f73a097 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:51Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:endpointPickerRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:number: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:matchLabels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPorts: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:38:51Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55380' [e2e-llm-inference-service] uid: 6a1402f0-1d9f-4727-9ad7-b1977c3987c7 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] appProtocol: http [e2e-llm-inference-service] endpointPickerRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service [e2e-llm-inference-service] port: [e2e-llm-inference-service] number: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPorts: [e2e-llm-inference-service] - number: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:51Z' [e2e-llm-inference-service] message: Referenced by an HTTPRoute accepted by the parentRef Gateway [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:51Z' [e2e-llm-inference-service] message: Referenced ExtensionRef resolved successfully [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: networking.istio.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] kind: AuthPolicy [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-policies [e2e-llm-inference-service] app.kubernetes.io/managed-by: odh-model-controller [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:rules: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:authentication: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:public: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:anonymous: {} [e2e-llm-inference-service] f:credentials: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:overrides: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fairness: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:response: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:success: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:headers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:x-gateway-inference-fairness-id: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:x-gateway-inference-objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:targetRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Accepted"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Enforced"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:38:38Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route-authn [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55187' [e2e-llm-inference-service] uid: 826b0aed-bd2c-49e2-8d33-c94b3ad3acc5 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] rules: [e2e-llm-inference-service] authentication: [e2e-llm-inference-service] public: [e2e-llm-inference-service] anonymous: {} [e2e-llm-inference-service] credentials: {} [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] overrides: [e2e-llm-inference-service] fairness: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] objective: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] response: [e2e-llm-inference-service] success: [e2e-llm-inference-service] headers: [e2e-llm-inference-service] x-gateway-inference-fairness-id: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.fairness [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] x-gateway-inference-objective: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.objective [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route [e2e-llm-inference-service] status: [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] message: AuthPolicy has been accepted [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:38Z' [e2e-llm-inference-service] message: AuthPolicy has been successfully enforced [e2e-llm-inference-service] reason: Enforced [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Enforced [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: leaderworkerset.x-k8s.io/v1 [e2e-llm-inference-service] kind: LeaderWorkerSet [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: leaderworkerset.x-k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:leaderWorkerTemplate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:leaderTemplate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:add: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:size: {} [e2e-llm-inference-service] f:workerTemplate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:add: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:rolloutStrategy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupPolicy: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] - apiVersion: leaderworkerset.x-k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:conditions: {} [e2e-llm-inference-service] f:hpaPodSelector: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] manager: lws [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:40:17Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '56619' [e2e-llm-inference-service] uid: 11af2961-0c70-4a78-8acb-69f5f7deb7b4 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] leaderWorkerTemplate: [e2e-llm-inference-service] leaderTemplate: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-leader [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] spec: [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] name: main [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /home [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - mountPath: /tmp [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - mountPath: /dev/shm [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - mountPath: /models [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - mountPath: /var/run/kserve/tls [e2e-llm-inference-service] name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] name: storage-initializer [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] serviceAccountName: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] restartPolicy: RecreateGroupOnPodRestart [e2e-llm-inference-service] size: 2 [e2e-llm-inference-service] workerTemplate: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] spec: [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - --served-model-name [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --enable-ssl-refresh [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] command: [e2e-llm-inference-service] - vllm [e2e-llm-inference-service] - serve [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] name: main [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] add: [e2e-llm-inference-service] - IPC_LOCK [e2e-llm-inference-service] - SYS_RAWIO [e2e-llm-inference-service] - NET_RAW [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /home [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - mountPath: /tmp [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - mountPath: /dev/shm [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - mountPath: /models [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - mountPath: /var/run/kserve/tls [e2e-llm-inference-service] name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] name: storage-initializer [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - mountPath: /mnt/models [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] serviceAccountName: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: home [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: tmp-dir [e2e-llm-inference-service] - emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 8Gi [e2e-llm-inference-service] name: dshm [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: model-cache [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] - emptyDir: {} [e2e-llm-inference-service] name: kserve-provision-location [e2e-llm-inference-service] networkConfig: [e2e-llm-inference-service] subdomainPolicy: Shared [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] rolloutStrategy: [e2e-llm-inference-service] rollingUpdateConfiguration: [e2e-llm-inference-service] maxSurge: 0 [e2e-llm-inference-service] maxUnavailable: 1 [e2e-llm-inference-service] partition: 0 [e2e-llm-inference-service] type: RollingUpdate [e2e-llm-inference-service] startupPolicy: LeaderCreated [e2e-llm-inference-service] status: [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:38:36Z' [e2e-llm-inference-service] message: Replicas are progressing [e2e-llm-inference-service] reason: GroupsProgressing [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: Progressing [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:40:17Z' [e2e-llm-inference-service] message: All replicas are ready [e2e-llm-inference-service] reason: AllGroupsReady [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Available [e2e-llm-inference-service] hpaPodSelector: leaderworkerset.sigs.k8s.io/name=llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn,leaderworkerset.sigs.k8s.io/worker-index=0 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55107' [e2e-llm-inference-service] uid: 3508fafd-5cab-48bf-a306-0f3eb351b73c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:52Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:52Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55399' [e2e-llm-inference-service] uid: a83f8e91-99c2-47d2-9266-8ee40ea42923 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-fb-opt-125m-route-dc21cb14-inference--ip-1de604e1.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55151' [e2e-llm-inference-service] uid: d828fcf0-1e10-47c6-825a-b5f832c4a089 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55107' [e2e-llm-inference-service] uid: 3508fafd-5cab-48bf-a306-0f3eb351b73c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:52Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:52Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55399' [e2e-llm-inference-service] uid: a83f8e91-99c2-47d2-9266-8ee40ea42923 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-fb-opt-125m-route-dc21cb14-inference--ip-1de604e1.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55151' [e2e-llm-inference-service] uid: d828fcf0-1e10-47c6-825a-b5f832c4a089 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55107' [e2e-llm-inference-service] uid: 3508fafd-5cab-48bf-a306-0f3eb351b73c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:52Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:52Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55399' [e2e-llm-inference-service] uid: a83f8e91-99c2-47d2-9266-8ee40ea42923 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-fb-opt-125m-route-dc21cb14-inference--ip-1de604e1.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55151' [e2e-llm-inference-service] uid: d828fcf0-1e10-47c6-825a-b5f832c4a089 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"3ab18621-3197-4628-8c4a-65662dce7a67"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:extensionRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:portNumber: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPortNumber: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:38:37Z' [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-inference-pool [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] uid: 3ab18621-3197-4628-8c4a-65662dce7a67 [e2e-llm-inference-service] resourceVersion: '55078' [e2e-llm-inference-service] uid: b0eb1d1e-0725-488d-812f-ecb0c699db83 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] extensionRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-epp-service [e2e-llm-inference-service] portNumber: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPortNumber: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parent: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '1970-01-01T00:00:00Z' [e2e-llm-inference-service] message: Waiting for controller [e2e-llm-inference-service] reason: Pending [e2e-llm-inference-service] status: Unknown [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Status [e2e-llm-inference-service] name: default [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:53:34Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-leader [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] apps.kubernetes.io/pod-index: '0' [e2e-llm-inference-service] controller-revision-hash: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-57dd9f69d [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: 269370d6a88f087dbcb80280bf4153dc63eb8499 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7f6c8cfc6 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '0' [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] statefulset.kubernetes.io/pod-name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] timestamp: '2026-07-02T13:53:10Z' [e2e-llm-inference-service] window: 14.893s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 101140938n [e2e-llm-inference-service] memory: 2400920Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:53:34Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload-worker [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] apps.kubernetes.io/pod-index: '1' [e2e-llm-inference-service] controller-revision-hash: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-5c59c86cbf [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-index: '0' [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/group-key: 269370d6a88f087dbcb80280bf4153dc63eb8499 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/template-revision-hash: 7f6c8cfc6 [e2e-llm-inference-service] leaderworkerset.sigs.k8s.io/worker-index: '1' [e2e-llm-inference-service] statefulset.kubernetes.io/pod-name: llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 [e2e-llm-inference-service] timestamp: '2026-07-02T13:53:16Z' [e2e-llm-inference-service] window: 13.267s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 4083213n [e2e-llm-inference-service] memory: 2410636Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:53:34Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-fb-opt-125m-route-dc21cb14 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 5688c7c666 [e2e-llm-inference-service] timestamp: '2026-07-02T13:53:09Z' [e2e-llm-inference-service] window: 11.578s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 58090430n [e2e-llm-inference-service] memory: 31916Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_inference_service] [2026-07-02T13:53:34.740709] end - ❌ 902.398s: Missing true conditions: {'WorkloadsReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:38:52Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:39:12Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'WorkerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:39Z', 'message': 'LWS is progressing', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] _____________ test_llm_with_lora_adapters[single-lora-adapter-hf] ______________ [e2e-llm-inference-service] [gw1] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), chunked = False [e2e-llm-inference-service] response_conn = [e2e-llm-inference-service] preload_content = False, decode_content = False, enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] # conn.request() calls http.client.*.request, not the method in [e2e-llm-inference-service] # urllib3.request. It also calls makefile (recv) on the socket. [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn.request( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] enforce_content_length=enforce_content_length, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # We are swallowing BrokenPipeError (errno.EPIPE) since the server is [e2e-llm-inference-service] # legitimately able to close the connection after sending a valid response. [e2e-llm-inference-service] # With this behaviour, the received response is still readable. [e2e-llm-inference-service] except BrokenPipeError: [e2e-llm-inference-service] pass [e2e-llm-inference-service] except OSError as e: [e2e-llm-inference-service] # MacOS/Linux [e2e-llm-inference-service] # EPROTOTYPE and ECONNRESET are needed on macOS [e2e-llm-inference-service] # https://erickt.github.io/blog/2014/11/19/adventures-in-debugging-a-potential-osx-kernel-bug/ [e2e-llm-inference-service] # Condition changed later to emit ECONNRESET instead of only EPROTOTYPE. [e2e-llm-inference-service] if e.errno != errno.EPROTOTYPE and e.errno != errno.ECONNRESET: [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset the timeout for the recv() on the socket [e2e-llm-inference-service] read_timeout = timeout_obj.read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn.is_closed: [e2e-llm-inference-service] # In Python 3 socket.py will catch EAGAIN and return None when you [e2e-llm-inference-service] # try and read into the file pointer created by http.client, which [e2e-llm-inference-service] # instead raises a BadStatusLine exception. Instead of catching [e2e-llm-inference-service] # the exception and assuming all BadStatusLine exceptions are read [e2e-llm-inference-service] # timeouts, check for a zero timeout before making the request. [e2e-llm-inference-service] if read_timeout == 0: [e2e-llm-inference-service] raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={read_timeout})" [e2e-llm-inference-service] ) [e2e-llm-inference-service] conn.timeout = read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] # Receive the response from the server [e2e-llm-inference-service] try: [e2e-llm-inference-service] > response = conn.getresponse() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:534: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def getresponse( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] ) -> HTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get the response from the server. [e2e-llm-inference-service] [e2e-llm-inference-service] If the HTTPConnection is in the correct state, returns an instance of HTTPResponse or of whatever object is returned by the response_class variable. [e2e-llm-inference-service] [e2e-llm-inference-service] If a request has not been sent or if a previous response has not be handled, ResponseNotReady is raised. If the HTTP response indicates that the connection should be closed, then it will be closed before the response is returned. When the connection is closed, the underlying socket is closed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Raise the same error as http.client.HTTPConnection [e2e-llm-inference-service] if self._response_options is None: [e2e-llm-inference-service] raise ResponseNotReady() [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset this attribute for being used again. [e2e-llm-inference-service] resp_options = self._response_options [e2e-llm-inference-service] self._response_options = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Since the connection's timeout value may have been updated [e2e-llm-inference-service] # we need to set the timeout on the socket. [e2e-llm-inference-service] self.sock.settimeout(self.timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] # This is needed here to avoid circular import errors [e2e-llm-inference-service] from .response import HTTPResponse [e2e-llm-inference-service] [e2e-llm-inference-service] # Save a reference to the shutdown function before ownership is passed [e2e-llm-inference-service] # to httplib_response [e2e-llm-inference-service] # TODO should we implement it everywhere? [e2e-llm-inference-service] _shutdown = getattr(self.sock, "shutdown", None) [e2e-llm-inference-service] [e2e-llm-inference-service] # Get the response from http.client.HTTPConnection [e2e-llm-inference-service] > httplib_response = super().getresponse() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:571: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def getresponse(self): [e2e-llm-inference-service] """Get the response from the server. [e2e-llm-inference-service] [e2e-llm-inference-service] If the HTTPConnection is in the correct state, returns an [e2e-llm-inference-service] instance of HTTPResponse or of whatever object is returned by [e2e-llm-inference-service] the response_class variable. [e2e-llm-inference-service] [e2e-llm-inference-service] If a request has not been sent or if a previous response has [e2e-llm-inference-service] not be handled, ResponseNotReady is raised. If the HTTP [e2e-llm-inference-service] response indicates that the connection should be closed, then [e2e-llm-inference-service] it will be closed before the response is returned. When the [e2e-llm-inference-service] connection is closed, the underlying socket is closed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] # if a prior response has been completed, then forget about it. [e2e-llm-inference-service] if self.__response and self.__response.isclosed(): [e2e-llm-inference-service] self.__response = None [e2e-llm-inference-service] [e2e-llm-inference-service] # if a prior response exists, then it must be completed (otherwise, we [e2e-llm-inference-service] # cannot read this response's header to determine the connection-close [e2e-llm-inference-service] # behavior) [e2e-llm-inference-service] # [e2e-llm-inference-service] # note: if a prior response existed, but was connection-close, then the [e2e-llm-inference-service] # socket and response were made independent of this HTTPConnection [e2e-llm-inference-service] # object since a new request requires that we open a whole new [e2e-llm-inference-service] # connection [e2e-llm-inference-service] # [e2e-llm-inference-service] # this means the prior response had one of two states: [e2e-llm-inference-service] # 1) will_close: this connection was reset and the prior socket and [e2e-llm-inference-service] # response operate independently [e2e-llm-inference-service] # 2) persistent: the response was retained and we await its [e2e-llm-inference-service] # isclosed() status to become true. [e2e-llm-inference-service] # [e2e-llm-inference-service] if self.__state != _CS_REQ_SENT or self.__response: [e2e-llm-inference-service] raise ResponseNotReady(self.__state) [e2e-llm-inference-service] [e2e-llm-inference-service] if self.debuglevel > 0: [e2e-llm-inference-service] response = self.response_class(self.sock, self.debuglevel, [e2e-llm-inference-service] method=self._method) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = self.response_class(self.sock, method=self._method) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > response.begin() [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:1395: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def begin(self): [e2e-llm-inference-service] if self.headers is not None: [e2e-llm-inference-service] # we've already started reading the response [e2e-llm-inference-service] return [e2e-llm-inference-service] [e2e-llm-inference-service] # read until we get a non-100 response [e2e-llm-inference-service] while True: [e2e-llm-inference-service] > version, status, reason = self._read_status() [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:325: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def _read_status(self): [e2e-llm-inference-service] > line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1") [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:286: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] b = [e2e-llm-inference-service] [e2e-llm-inference-service] def readinto(self, b): [e2e-llm-inference-service] """Read up to len(b) bytes into the writable buffer *b* and return [e2e-llm-inference-service] the number of bytes read. If the socket is non-blocking and no bytes [e2e-llm-inference-service] are available, None is returned. [e2e-llm-inference-service] [e2e-llm-inference-service] If *b* is non-empty, a 0 return value indicates that the connection [e2e-llm-inference-service] was shutdown at the other end. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self._checkClosed() [e2e-llm-inference-service] self._checkReadable() [e2e-llm-inference-service] if self._timeout_occurred: [e2e-llm-inference-service] raise OSError("cannot read from timed out object") [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return self._sock.recv_into(b) [e2e-llm-inference-service] E TimeoutError: timed out [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/socket.py:718: TimeoutError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] > response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:787: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), chunked = False [e2e-llm-inference-service] response_conn = [e2e-llm-inference-service] preload_content = False, decode_content = False, enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] # conn.request() calls http.client.*.request, not the method in [e2e-llm-inference-service] # urllib3.request. It also calls makefile (recv) on the socket. [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn.request( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] enforce_content_length=enforce_content_length, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # We are swallowing BrokenPipeError (errno.EPIPE) since the server is [e2e-llm-inference-service] # legitimately able to close the connection after sending a valid response. [e2e-llm-inference-service] # With this behaviour, the received response is still readable. [e2e-llm-inference-service] except BrokenPipeError: [e2e-llm-inference-service] pass [e2e-llm-inference-service] except OSError as e: [e2e-llm-inference-service] # MacOS/Linux [e2e-llm-inference-service] # EPROTOTYPE and ECONNRESET are needed on macOS [e2e-llm-inference-service] # https://erickt.github.io/blog/2014/11/19/adventures-in-debugging-a-potential-osx-kernel-bug/ [e2e-llm-inference-service] # Condition changed later to emit ECONNRESET instead of only EPROTOTYPE. [e2e-llm-inference-service] if e.errno != errno.EPROTOTYPE and e.errno != errno.ECONNRESET: [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset the timeout for the recv() on the socket [e2e-llm-inference-service] read_timeout = timeout_obj.read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn.is_closed: [e2e-llm-inference-service] # In Python 3 socket.py will catch EAGAIN and return None when you [e2e-llm-inference-service] # try and read into the file pointer created by http.client, which [e2e-llm-inference-service] # instead raises a BadStatusLine exception. Instead of catching [e2e-llm-inference-service] # the exception and assuming all BadStatusLine exceptions are read [e2e-llm-inference-service] # timeouts, check for a zero timeout before making the request. [e2e-llm-inference-service] if read_timeout == 0: [e2e-llm-inference-service] raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={read_timeout})" [e2e-llm-inference-service] ) [e2e-llm-inference-service] conn.timeout = read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] # Receive the response from the server [e2e-llm-inference-service] try: [e2e-llm-inference-service] response = conn.getresponse() [e2e-llm-inference-service] except (BaseSSLError, OSError) as e: [e2e-llm-inference-service] > self._raise_timeout(err=e, url=url, timeout_value=read_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:536: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] err = TimeoutError('timed out') [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] timeout_value = 60 [e2e-llm-inference-service] [e2e-llm-inference-service] def _raise_timeout( [e2e-llm-inference-service] self, [e2e-llm-inference-service] err: BaseSSLError | OSError | SocketTimeout, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] timeout_value: _TYPE_TIMEOUT | None, [e2e-llm-inference-service] ) -> None: [e2e-llm-inference-service] """Is the error actually a timeout? Will raise a ReadTimeout or pass""" [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(err, SocketTimeout): [e2e-llm-inference-service] > raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={timeout_value})" [e2e-llm-inference-service] ) from err [e2e-llm-inference-service] E urllib3.exceptions.ReadTimeoutError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:367: ReadTimeoutError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = , stream = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), verify = '/tmp/ca.crt' [e2e-llm-inference-service] cert = None, proxies = OrderedDict() [e2e-llm-inference-service] [e2e-llm-inference-service] def send( [e2e-llm-inference-service] self, request, stream=False, timeout=None, verify=True, cert=None, proxies=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Sends PreparedRequest object. Returns Response object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param request: The :class:`PreparedRequest ` being sent. [e2e-llm-inference-service] :param stream: (optional) Whether to stream the request content. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple or urllib3 Timeout object [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether [e2e-llm-inference-service] we verify the server's TLS certificate, or a string, in which case it [e2e-llm-inference-service] must be a path to a CA bundle to use [e2e-llm-inference-service] :param cert: (optional) Any user-provided SSL certificate to be trusted. [e2e-llm-inference-service] :param proxies: (optional) The proxies dictionary to apply to the request. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn = self.get_connection_with_tls_context( [e2e-llm-inference-service] request, verify, proxies=proxies, cert=cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] except LocationValueError as e: [e2e-llm-inference-service] raise InvalidURL(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] self.cert_verify(conn, request.url, verify, cert) [e2e-llm-inference-service] url = self.request_url(request, proxies) [e2e-llm-inference-service] self.add_headers( [e2e-llm-inference-service] request, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] verify=verify, [e2e-llm-inference-service] cert=cert, [e2e-llm-inference-service] proxies=proxies, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] chunked = not (request.body is None or "Content-Length" in request.headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(timeout, tuple): [e2e-llm-inference-service] try: [e2e-llm-inference-service] connect, read = timeout [e2e-llm-inference-service] timeout = TimeoutSauce(connect=connect, read=read) [e2e-llm-inference-service] except ValueError: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, " [e2e-llm-inference-service] f"or a single float to set both timeouts to the same value." [e2e-llm-inference-service] ) [e2e-llm-inference-service] elif isinstance(timeout, TimeoutSauce): [e2e-llm-inference-service] pass [e2e-llm-inference-service] else: [e2e-llm-inference-service] timeout = TimeoutSauce(connect=timeout, read=timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > resp = conn.urlopen( [e2e-llm-inference-service] method=request.method, [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] body=request.body, [e2e-llm-inference-service] headers=request.headers, [e2e-llm-inference-service] redirect=False, [e2e-llm-inference-service] assert_same_host=False, [e2e-llm-inference-service] preload_content=False, [e2e-llm-inference-service] decode_content=False, [e2e-llm-inference-service] retries=self.max_retries, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/adapters.py:667: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=7, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=6, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=5, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=4, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=3, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=2, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = RemoteDisconnected('Remote end closed connection without response') [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=1, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "What is Kubernetes?", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '82', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] > retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:841: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] response = None [e2e-llm-inference-service] error = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] _pool = [e2e-llm-inference-service] _stacktrace = [e2e-llm-inference-service] [e2e-llm-inference-service] def increment( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str | None = None, [e2e-llm-inference-service] url: str | None = None, [e2e-llm-inference-service] response: BaseHTTPResponse | None = None, [e2e-llm-inference-service] error: Exception | None = None, [e2e-llm-inference-service] _pool: ConnectionPool | None = None, [e2e-llm-inference-service] _stacktrace: TracebackType | None = None, [e2e-llm-inference-service] ) -> Self: [e2e-llm-inference-service] """Return a new Retry object with incremented retry counters. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response: A response object, or None, if the server did not [e2e-llm-inference-service] return a response. [e2e-llm-inference-service] :type response: :class:`~urllib3.response.BaseHTTPResponse` [e2e-llm-inference-service] :param Exception error: An error encountered during the request, or [e2e-llm-inference-service] None if the response was received successfully. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: A new ``Retry`` object. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if self.total is False and error: [e2e-llm-inference-service] # Disabled, indicate to re-raise the error. [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] [e2e-llm-inference-service] total = self.total [e2e-llm-inference-service] if total is not None: [e2e-llm-inference-service] total -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] connect = self.connect [e2e-llm-inference-service] read = self.read [e2e-llm-inference-service] redirect = self.redirect [e2e-llm-inference-service] status_count = self.status [e2e-llm-inference-service] other = self.other [e2e-llm-inference-service] cause = "unknown" [e2e-llm-inference-service] status = None [e2e-llm-inference-service] redirect_location = None [e2e-llm-inference-service] [e2e-llm-inference-service] if error and self._is_connection_error(error): [e2e-llm-inference-service] # Connect retry? [e2e-llm-inference-service] if connect is False: [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif connect is not None: [e2e-llm-inference-service] connect -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error and self._is_read_error(error): [e2e-llm-inference-service] # Read retry? [e2e-llm-inference-service] if read is False or method is None or not self._is_method_retryable(method): [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif read is not None: [e2e-llm-inference-service] read -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error: [e2e-llm-inference-service] # Other retry? [e2e-llm-inference-service] if other is not None: [e2e-llm-inference-service] other -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif response and response.get_redirect_location(): [e2e-llm-inference-service] # Redirect retry? [e2e-llm-inference-service] if redirect is not None: [e2e-llm-inference-service] redirect -= 1 [e2e-llm-inference-service] cause = "too many redirects" [e2e-llm-inference-service] response_redirect_location = response.get_redirect_location() [e2e-llm-inference-service] if response_redirect_location: [e2e-llm-inference-service] redirect_location = response_redirect_location [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Incrementing because of a server error like a 500 in [e2e-llm-inference-service] # status_forcelist and the given method is in the allowed_methods [e2e-llm-inference-service] cause = ResponseError.GENERIC_ERROR [e2e-llm-inference-service] if response and response.status: [e2e-llm-inference-service] if status_count is not None: [e2e-llm-inference-service] status_count -= 1 [e2e-llm-inference-service] cause = ResponseError.SPECIFIC_ERROR.format(status_code=response.status) [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] history = self.history + ( [e2e-llm-inference-service] RequestHistory(method, url, error, status, redirect_location), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] new_retry = self.new( [e2e-llm-inference-service] total=total, [e2e-llm-inference-service] connect=connect, [e2e-llm-inference-service] read=read, [e2e-llm-inference-service] redirect=redirect, [e2e-llm-inference-service] status=status_count, [e2e-llm-inference-service] other=other, [e2e-llm-inference-service] history=history, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if new_retry.is_exhausted(): [e2e-llm-inference-service] reason = error or ResponseError(cause) [e2e-llm-inference-service] > raise MaxRetryError(_pool, url, reason) from reason # type: ignore[arg-type] [e2e-llm-inference-service] E urllib3.exceptions.MaxRetryError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/util/retry.py:519: MaxRetryError [e2e-llm-inference-service] [e2e-llm-inference-service] During handling of the above exception, another exception occurred: [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = LoRATestCase(base_refs=['router-no-scheduler', 'workload-single-cpu', 'model-fb-opt-125m-with-lora-hf'], prompt='What ...ice_name='lora-single-adapter-test', endpoint='/v1/completions', max_tokens=100, wait_timeout=900, response_timeout=60) [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] LoRATestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-no-scheduler", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="What is Kubernetes?", [e2e-llm-inference-service] expected_adapter_names=["lora-adapter-1"], [e2e-llm-inference-service] service_name="lora-single-adapter-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.llminferenceservice, [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] id="single-lora-adapter-hf", [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] LoRATestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-no-scheduler", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-multiple-lora", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="Explain machine learning in simple terms.", [e2e-llm-inference-service] expected_adapter_names=["lora-adapter-1", "lora-adapter-2"], [e2e-llm-inference-service] service_name="lora-multiple-adapters-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.llminferenceservice, [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] id="multiple-lora-adapters", [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_with_lora_adapters(test_case: LoRATestCase): [e2e-llm-inference-service] """Test LLMInferenceService with LoRA adapters.""" [e2e-llm-inference-service] > run_lora_test(test_case) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_lora_adapters.py:248: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = LoRATestCase(base_refs=['router-no-scheduler', 'workload-single-cpu', 'model-fb-opt-125m-with-lora-hf'], prompt='What ...ice_name='lora-single-adapter-test', endpoint='/v1/completions', max_tokens=100, wait_timeout=900, response_timeout=60) [e2e-llm-inference-service] [e2e-llm-inference-service] def run_lora_test(test_case: LoRATestCase): [e2e-llm-inference-service] """Execute a LoRA adapter test case.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] test_id = generate_test_id(test_case) [e2e-llm-inference-service] service_name = test_case.service_name or f"lora-test-{test_id}" [e2e-llm-inference-service] model_name = _get_model_name_from_configs(test_case.base_refs) [e2e-llm-inference-service] created_configs = [] [e2e-llm-inference-service] llm_service = None [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Create unique LLMInferenceServiceConfig resources for each base ref [e2e-llm-inference-service] unique_base_refs = [] [e2e-llm-inference-service] for base_ref in test_case.base_refs: [e2e-llm-inference-service] if base_ref not in LLMINFERENCESERVICE_CONFIGS: [e2e-llm-inference-service] raise ValueError(f"Unknown base reference: {base_ref}") [e2e-llm-inference-service] [e2e-llm-inference-service] # Generate unique config name to avoid conflicts in parallel test runs [e2e-llm-inference-service] unique_config_name = generate_k8s_safe_suffix(base_ref, [service_name]) [e2e-llm-inference-service] unique_base_refs.append(unique_config_name) [e2e-llm-inference-service] [e2e-llm-inference-service] config = LLMINFERENCESERVICE_CONFIGS[base_ref] [e2e-llm-inference-service] config_body = { [e2e-llm-inference-service] "apiVersion": "serving.kserve.io/v1alpha1", [e2e-llm-inference-service] "kind": "LLMInferenceServiceConfig", [e2e-llm-inference-service] "metadata": { [e2e-llm-inference-service] "name": unique_config_name, [e2e-llm-inference-service] "namespace": KSERVE_TEST_NAMESPACE, [e2e-llm-inference-service] }, [e2e-llm-inference-service] "spec": config, [e2e-llm-inference-service] } [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info("Creating LLMInferenceServiceConfig: %s", unique_config_name) [e2e-llm-inference-service] _create_or_update_llmisvc_config( [e2e-llm-inference-service] kserve_client, config_body, KSERVE_TEST_NAMESPACE [e2e-llm-inference-service] ) [e2e-llm-inference-service] created_configs.append(unique_config_name) [e2e-llm-inference-service] [e2e-llm-inference-service] # Create the service with unique base refs [e2e-llm-inference-service] llm_service = build_llm_service_from_refs(service_name, unique_base_refs) [e2e-llm-inference-service] if not llm_service.metadata.annotations: [e2e-llm-inference-service] llm_service.metadata.annotations = {} [e2e-llm-inference-service] llm_service.metadata.annotations["security.opendatahub.io/enable-auth"] = ( [e2e-llm-inference-service] "false" [e2e-llm-inference-service] ) [e2e-llm-inference-service] logger.info("Creating LLMInferenceService: %s", service_name) [e2e-llm-inference-service] create_llmisvc(kserve_client, llm_service) [e2e-llm-inference-service] [e2e-llm-inference-service] # Wait for service to be ready [e2e-llm-inference-service] logger.info("Waiting for service %s to be ready...", service_name) [e2e-llm-inference-service] wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, llm_service, timeout_seconds=test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Get inference URL [e2e-llm-inference-service] base_url = get_llm_service_url(kserve_client, llm_service) [e2e-llm-inference-service] inference_url = base_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] # Test base model inference [e2e-llm-inference-service] logger.info("Testing base model inference...") [e2e-llm-inference-service] base_payload = { [e2e-llm-inference-service] "model": model_name, [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] [e2e-llm-inference-service] > base_response = post_with_retry( [e2e-llm-inference-service] inference_url, [e2e-llm-inference-service] json_data=base_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_lora_adapters.py:150: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] [e2e-llm-inference-service] def post_with_retry( [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] *, [e2e-llm-inference-service] headers: Dict = None, [e2e-llm-inference-service] json_data: Union[Dict, List] = None, [e2e-llm-inference-service] data: Union[str, bytes] = None, [e2e-llm-inference-service] stream: bool = False, [e2e-llm-inference-service] timeout: float = None, [e2e-llm-inference-service] total_retries: int = DEFAULT_RETRY_TOTAL, [e2e-llm-inference-service] backoff_factor: float = DEFAULT_RETRY_BACKOFF_FACTOR, [e2e-llm-inference-service] retry_status_codes=DEFAULT_RETRY_STATUS_CODES, [e2e-llm-inference-service] ) -> requests.Response: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Send POST request with retries for transient HTTP and network failures. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if json_data is not None and data is not None: [e2e-llm-inference-service] raise ValueError("Only one of json_data or data can be provided.") [e2e-llm-inference-service] [e2e-llm-inference-service] with _retry_session( [e2e-llm-inference-service] ["POST"], total_retries, backoff_factor, retry_status_codes [e2e-llm-inference-service] ) as session: [e2e-llm-inference-service] > return session.post( [e2e-llm-inference-service] url, [e2e-llm-inference-service] json=json_data, [e2e-llm-inference-service] data=data, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] common/http_retry.py:70: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] data = None [e2e-llm-inference-service] json = {'max_tokens': 100, 'model': 'facebook/opt-125m', 'prompt': 'What is Kubernetes?'} [e2e-llm-inference-service] kwargs = {'headers': None, 'stream': False, 'timeout': 60} [e2e-llm-inference-service] [e2e-llm-inference-service] def post(self, url, data=None, json=None, **kwargs): [e2e-llm-inference-service] r"""Sends a POST request. Returns :class:`Response` object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: URL for the new :class:`Request` object. [e2e-llm-inference-service] :param data: (optional) Dictionary, list of tuples, bytes, or file-like [e2e-llm-inference-service] object to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param json: (optional) json to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param \*\*kwargs: Optional arguments that ``request`` takes. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.request("POST", url, data=data, json=json, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:637: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = , method = 'POST' [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/lora-single-adapter-test/v1/completions' [e2e-llm-inference-service] params = None, data = None, headers = None, cookies = None, files = None [e2e-llm-inference-service] auth = None, timeout = 60, allow_redirects = True, proxies = {}, hooks = None [e2e-llm-inference-service] stream = False, verify = None, cert = None [e2e-llm-inference-service] json = {'max_tokens': 100, 'model': 'facebook/opt-125m', 'prompt': 'What is Kubernetes?'} [e2e-llm-inference-service] [e2e-llm-inference-service] def request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] params=None, [e2e-llm-inference-service] data=None, [e2e-llm-inference-service] headers=None, [e2e-llm-inference-service] cookies=None, [e2e-llm-inference-service] files=None, [e2e-llm-inference-service] auth=None, [e2e-llm-inference-service] timeout=None, [e2e-llm-inference-service] allow_redirects=True, [e2e-llm-inference-service] proxies=None, [e2e-llm-inference-service] hooks=None, [e2e-llm-inference-service] stream=None, [e2e-llm-inference-service] verify=None, [e2e-llm-inference-service] cert=None, [e2e-llm-inference-service] json=None, [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Constructs a :class:`Request `, prepares it and sends it. [e2e-llm-inference-service] Returns :class:`Response ` object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: method for the new :class:`Request` object. [e2e-llm-inference-service] :param url: URL for the new :class:`Request` object. [e2e-llm-inference-service] :param params: (optional) Dictionary or bytes to be sent in the query [e2e-llm-inference-service] string for the :class:`Request`. [e2e-llm-inference-service] :param data: (optional) Dictionary, list of tuples, bytes, or file-like [e2e-llm-inference-service] object to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param json: (optional) json to send in the body of the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param headers: (optional) Dictionary of HTTP Headers to send with the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param cookies: (optional) Dict or CookieJar object to send with the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param files: (optional) Dictionary of ``'filename': file-like-objects`` [e2e-llm-inference-service] for multipart encoding upload. [e2e-llm-inference-service] :param auth: (optional) Auth tuple or callable to enable [e2e-llm-inference-service] Basic/Digest/Custom HTTP Auth. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple [e2e-llm-inference-service] :param allow_redirects: (optional) Set to True by default. [e2e-llm-inference-service] :type allow_redirects: bool [e2e-llm-inference-service] :param proxies: (optional) Dictionary mapping protocol or protocol and [e2e-llm-inference-service] hostname to the URL of the proxy. [e2e-llm-inference-service] :param hooks: (optional) Dictionary mapping hook name to one event or [e2e-llm-inference-service] list of events, event must be callable. [e2e-llm-inference-service] :param stream: (optional) whether to immediately download the response [e2e-llm-inference-service] content. Defaults to ``False``. [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether we verify [e2e-llm-inference-service] the server's TLS certificate, or a string, in which case it must be a path [e2e-llm-inference-service] to a CA bundle to use. Defaults to ``True``. When set to [e2e-llm-inference-service] ``False``, requests will accept any TLS certificate presented by [e2e-llm-inference-service] the server, and will ignore hostname mismatches and/or expired [e2e-llm-inference-service] certificates, which will make your application vulnerable to [e2e-llm-inference-service] man-in-the-middle (MitM) attacks. Setting verify to ``False`` [e2e-llm-inference-service] may be useful during local development or testing. [e2e-llm-inference-service] :param cert: (optional) if String, path to ssl client cert file (.pem). [e2e-llm-inference-service] If Tuple, ('cert', 'key') pair. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Create the Request. [e2e-llm-inference-service] req = Request( [e2e-llm-inference-service] method=method.upper(), [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] files=files, [e2e-llm-inference-service] data=data or {}, [e2e-llm-inference-service] json=json, [e2e-llm-inference-service] params=params or {}, [e2e-llm-inference-service] auth=auth, [e2e-llm-inference-service] cookies=cookies, [e2e-llm-inference-service] hooks=hooks, [e2e-llm-inference-service] ) [e2e-llm-inference-service] prep = self.prepare_request(req) [e2e-llm-inference-service] [e2e-llm-inference-service] proxies = proxies or {} [e2e-llm-inference-service] [e2e-llm-inference-service] settings = self.merge_environment_settings( [e2e-llm-inference-service] prep.url, proxies, stream, verify, cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Send the request. [e2e-llm-inference-service] send_kwargs = { [e2e-llm-inference-service] "timeout": timeout, [e2e-llm-inference-service] "allow_redirects": allow_redirects, [e2e-llm-inference-service] } [e2e-llm-inference-service] send_kwargs.update(settings) [e2e-llm-inference-service] > resp = self.send(prep, **send_kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:589: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = [e2e-llm-inference-service] kwargs = {'cert': None, 'proxies': OrderedDict(), 'stream': False, 'timeout': 60, ...} [e2e-llm-inference-service] allow_redirects = True, stream = False, hooks = {'response': []} [e2e-llm-inference-service] adapter = [e2e-llm-inference-service] start = 1782999572.0615048 [e2e-llm-inference-service] [e2e-llm-inference-service] def send(self, request, **kwargs): [e2e-llm-inference-service] """Send a given PreparedRequest. [e2e-llm-inference-service] [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Set defaults that the hooks can utilize to ensure they always have [e2e-llm-inference-service] # the correct parameters to reproduce the previous request. [e2e-llm-inference-service] kwargs.setdefault("stream", self.stream) [e2e-llm-inference-service] kwargs.setdefault("verify", self.verify) [e2e-llm-inference-service] kwargs.setdefault("cert", self.cert) [e2e-llm-inference-service] if "proxies" not in kwargs: [e2e-llm-inference-service] kwargs["proxies"] = resolve_proxies(request, self.proxies, self.trust_env) [e2e-llm-inference-service] [e2e-llm-inference-service] # It's possible that users might accidentally send a Request object. [e2e-llm-inference-service] # Guard against that specific failure case. [e2e-llm-inference-service] if isinstance(request, Request): [e2e-llm-inference-service] raise ValueError("You can only send PreparedRequests.") [e2e-llm-inference-service] [e2e-llm-inference-service] # Set up variables needed for resolve_redirects and dispatching of hooks [e2e-llm-inference-service] allow_redirects = kwargs.pop("allow_redirects", True) [e2e-llm-inference-service] stream = kwargs.get("stream") [e2e-llm-inference-service] hooks = request.hooks [e2e-llm-inference-service] [e2e-llm-inference-service] # Get the appropriate adapter to use [e2e-llm-inference-service] adapter = self.get_adapter(url=request.url) [e2e-llm-inference-service] [e2e-llm-inference-service] # Start time (approximately) of the request [e2e-llm-inference-service] start = preferred_clock() [e2e-llm-inference-service] [e2e-llm-inference-service] # Send the request [e2e-llm-inference-service] > r = adapter.send(request, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:703: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = , stream = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), verify = '/tmp/ca.crt' [e2e-llm-inference-service] cert = None, proxies = OrderedDict() [e2e-llm-inference-service] [e2e-llm-inference-service] def send( [e2e-llm-inference-service] self, request, stream=False, timeout=None, verify=True, cert=None, proxies=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Sends PreparedRequest object. Returns Response object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param request: The :class:`PreparedRequest ` being sent. [e2e-llm-inference-service] :param stream: (optional) Whether to stream the request content. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple or urllib3 Timeout object [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether [e2e-llm-inference-service] we verify the server's TLS certificate, or a string, in which case it [e2e-llm-inference-service] must be a path to a CA bundle to use [e2e-llm-inference-service] :param cert: (optional) Any user-provided SSL certificate to be trusted. [e2e-llm-inference-service] :param proxies: (optional) The proxies dictionary to apply to the request. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn = self.get_connection_with_tls_context( [e2e-llm-inference-service] request, verify, proxies=proxies, cert=cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] except LocationValueError as e: [e2e-llm-inference-service] raise InvalidURL(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] self.cert_verify(conn, request.url, verify, cert) [e2e-llm-inference-service] url = self.request_url(request, proxies) [e2e-llm-inference-service] self.add_headers( [e2e-llm-inference-service] request, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] verify=verify, [e2e-llm-inference-service] cert=cert, [e2e-llm-inference-service] proxies=proxies, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] chunked = not (request.body is None or "Content-Length" in request.headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(timeout, tuple): [e2e-llm-inference-service] try: [e2e-llm-inference-service] connect, read = timeout [e2e-llm-inference-service] timeout = TimeoutSauce(connect=connect, read=read) [e2e-llm-inference-service] except ValueError: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, " [e2e-llm-inference-service] f"or a single float to set both timeouts to the same value." [e2e-llm-inference-service] ) [e2e-llm-inference-service] elif isinstance(timeout, TimeoutSauce): [e2e-llm-inference-service] pass [e2e-llm-inference-service] else: [e2e-llm-inference-service] timeout = TimeoutSauce(connect=timeout, read=timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] resp = conn.urlopen( [e2e-llm-inference-service] method=request.method, [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] body=request.body, [e2e-llm-inference-service] headers=request.headers, [e2e-llm-inference-service] redirect=False, [e2e-llm-inference-service] assert_same_host=False, [e2e-llm-inference-service] preload_content=False, [e2e-llm-inference-service] decode_content=False, [e2e-llm-inference-service] retries=self.max_retries, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] except (ProtocolError, OSError) as err: [e2e-llm-inference-service] raise ConnectionError(err, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] except MaxRetryError as e: [e2e-llm-inference-service] if isinstance(e.reason, ConnectTimeoutError): [e2e-llm-inference-service] # TODO: Remove this in 3.0.0: see #2811 [e2e-llm-inference-service] if not isinstance(e.reason, NewConnectionError): [e2e-llm-inference-service] raise ConnectTimeout(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, ResponseError): [e2e-llm-inference-service] raise RetryError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, _ProxyError): [e2e-llm-inference-service] raise ProxyError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, _SSLError): [e2e-llm-inference-service] # This branch is for urllib3 v1.22 and later. [e2e-llm-inference-service] raise SSLError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] > raise ConnectionError(e, request=request) [e2e-llm-inference-service] E requests.exceptions.ConnectionError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/adapters.py:700: ConnectionError [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_with_lora_adapters] [2026-07-02T13:37:55.549938] start - args=(), kwargs={'test_case': LoRATestCase(base_refs=['router-no-scheduler', 'workload-single-cpu', 'model-fb-opt-125m-with-lora-hf'], prompt='What is Kubernetes?', expected_adapter_names=['lora-adapter-1'], service_name='lora-single-adapter-test', endpoint='/v1/completions', max_tokens=100, wait_timeout=900, response_timeout=60)} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:116 Creating LLMInferenceServiceConfig: router-no-scheduler-lora-single-3cece2ef [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig router-no-scheduler-lora-single-3cece2ef in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig router-no-scheduler-lora-single-3cece2ef [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig router-no-scheduler-lora-single-3cece2ef [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:116 Creating LLMInferenceServiceConfig: workload-single-cpu-lora-single-6bc19feb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig workload-single-cpu-lora-single-6bc19feb in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig workload-single-cpu-lora-single-6bc19feb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig workload-single-cpu-lora-single-6bc19feb [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:116 Creating LLMInferenceServiceConfig: model-fb-opt-125m-with-lora-hf-721eb685 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig model-fb-opt-125m-with-lora-hf-721eb685 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig model-fb-opt-125m-with-lora-hf-721eb685 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig model-fb-opt-125m-with-lora-hf-721eb685 [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:129 Creating LLMInferenceService: lora-single-adapter-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-02T13:37:55.699828] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'lora-single-adapter-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-no-scheduler-lora-single-3cece2ef'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-lora-single-6bc19feb'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-with-lora-hf-721eb685'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-02T13:37:55.791968] end - ✅ in 0.092s [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:133 Waiting for service lora-single-adapter-test to be ready... [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-02T13:37:55.792147] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'lora-single-adapter-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-no-scheduler-lora-single-3cece2ef'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-lora-single-6bc19feb'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-with-lora-hf-721eb685'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={'timeout_seconds': 900} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-02T13:38:01Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/lora-single-adapter-test-kserve-route: v1.HTTPRouteStatus{RouteStatus:v1.RouteStatus{Parents:[]v1.RouteParentStatus(nil)}}]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:38:01Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:01Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:38:01Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/lora-single-adapter-test-kserve-route: v1.HTTPRouteStatus{RouteStatus:v1.RouteStatus{Parents:[]v1.RouteParentStatus(nil)}}]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:38:01Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/lora-single-adapter-test-kserve-route: v1.HTTPRouteStatus{RouteStatus:v1.RouteStatus{Parents:[]v1.RouteParentStatus(nil)}}]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:38:01Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-02T13:38:11Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:38:01Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:38:01Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:38:11Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:38:11Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:38:01Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [wait_for_llm_isvc_ready] [2026-07-02T13:39:32.052730] end - ✅ in 96.260s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [get_llm_service_url] [2026-07-02T13:39:32.052864] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'lora-single-adapter-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-no-scheduler-lora-single-3cece2ef'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-lora-single-6bc19feb'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-with-lora-hf-721eb685'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [get_llm_service_url] [2026-07-02T13:39:32.060644] end - ✅ in 0.008s [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:143 Testing base model inference... [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=7, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=6, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=5, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=4, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=3, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=2, connect=None, read=None, redirect=None, status=None)) after connection broken by 'RemoteDisconnected('Remote end closed connection without response')': /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=1, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=0, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:181 Cleaning up service lora-single-adapter-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [delete_llmisvc] [2026-07-02T13:54:36.673483] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'lora-single-adapter-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-no-scheduler-lora-single-3cece2ef'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-lora-single-6bc19feb'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-with-lora-hf-721eb685'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for lora-single-adapter-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_L [e2e-llm-inference-service] OGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:termi [e2e-llm-inference-service] nationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n [e2e-llm-inference-service] 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n [e2e-llm-inference-service] 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 'resource_version': '55968',\n 'self_link' [e2e-llm-inference-service] : None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=(" [e2e-llm-inference-service] $hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find a [e2e-llm-inference-service] ll RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gi [e2e-llm-inference-service] d_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n [e2e-llm-inference-service] 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None}, [e2e-llm-inference-service] \n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile [e2e-llm-inference-service] ': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333 [e2e-llm-inference-service] ',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': No [e2e-llm-inference-service] ne,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_se [e2e-llm-inference-service] conds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'g [e2e-llm-inference-service] it_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, [e2e-llm-inference-service] 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n [e2e-llm-inference-service] 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlo [e2e-llm-inference-service] cal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n [e2e-llm-inference-service] 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LO [e2e-llm-inference-service] GGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:termin [e2e-llm-inference-service] ationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n [e2e-llm-inference-service] 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n [e2e-llm-inference-service] 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 'resource_version': '55968',\n 'self_link': [e2e-llm-inference-service] None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$ [e2e-llm-inference-service] hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find al [e2e-llm-inference-service] l RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid [e2e-llm-inference-service] _index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n [e2e-llm-inference-service] 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\ [e2e-llm-inference-service] n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile' [e2e-llm-inference-service] : None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333' [e2e-llm-inference-service] ,\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': Non [e2e-llm-inference-service] e,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_sec [e2e-llm-inference-service] onds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'gi [e2e-llm-inference-service] t_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, ' [e2e-llm-inference-service] size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n [e2e-llm-inference-service] 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzloc [e2e-llm-inference-service] al()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n [e2e-llm-inference-service] 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha [e2e-llm-inference-service] .kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n [e2e-llm-inference-service] 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n [e2e-llm-inference-service] 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n [e2e-llm-inference-service] 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n [e2e-llm-inference-service] 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n [e2e-llm-inference-service] 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n [e2e-llm-inference-service] 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n [e2e-llm-inference-service] 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 'resource_version': '55968',\n 'self_link': None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n ' for hca_dir in '\n [e2e-llm-inference-service] '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'echo "${active_hcas[*]}")\n'\n ' '\n [e2e-llm-inference-service] 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n'\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if g [e2e-llm-inference-service] rep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n ' fi\n'\n ' d [e2e-llm-inference-service] one\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common: '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n [e2e-llm-inference-service] ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed '\n [e2e-llm-inference-service] 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-rele [e2e-llm-inference-service] ase-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n [e2e-llm-inference-service] 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds', [e2e-llm-inference-service] \n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n [e2e-llm-inference-service] 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n [e2e-llm-inference-service] 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n [e2e-llm-inference-service] 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n [e2e-llm-inference-service] 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n [e2e-llm-inference-service] 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n [e2e-llm-inference-service] 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n [e2e-llm-inference-service] 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'reso [e2e-llm-inference-service] urces': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n [e2e-llm-inference-service] 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n [e2e-llm-inference-service] 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '68847',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for lora-single-adapter-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 36, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit' [e2e-llm-inference-service] : {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 54, 36, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 'r [e2e-llm-inference-service] esource_version': '68862',\n 'self_link': None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n [e2e-llm-inference-service] 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n [e2e-llm-inference-service] ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n [e2e-llm-inference-service] 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-acces [e2e-llm-inference-service] s-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n [e2e-llm-inference-service] 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n [e2e-llm-inference-service] 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n [e2e-llm-inference-service] 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_prob [e2e-llm-inference-service] e': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': [e2e-llm-inference-service] 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'g [e2e-llm-inference-service] ce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n [e2e-llm-inference-service] 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n [e2e-llm-inference-service] 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': [e2e-llm-inference-service] datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n [e2e-llm-inference-service] 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n [e2e-llm-inference-service] 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 36, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': [e2e-llm-inference-service] {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 54, 36, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 're [e2e-llm-inference-service] source_version': '68862',\n 'self_link': None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n [e2e-llm-inference-service] 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n [e2e-llm-inference-service] ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n [e2e-llm-inference-service] 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access [e2e-llm-inference-service] -log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n [e2e-llm-inference-service] 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n [e2e-llm-inference-service] 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n [e2e-llm-inference-service] 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe [e2e-llm-inference-service] ': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': [e2e-llm-inference-service] 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gc [e2e-llm-inference-service] e_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n [e2e-llm-inference-service] 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n [e2e-llm-inference-service] 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': d [e2e-llm-inference-service] atetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n [e2e-llm-inference-service] 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n [e2e-llm-inference-service] 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n [e2e-llm-inference-service] 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 36, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n [e2e-llm-inference-service] 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n [e2e-llm-inference-service] 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n [e2e-llm-inference-service] 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {}, [e2e-llm-inference-service] \n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n [e2e-llm-inference-service] 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n [e2e-llm-inference-service] 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n [e2e-llm-inference-service] 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 54, 36, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 'resource_version': '68862',\n 'self_link': None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n [e2e-llm-inference-service] ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'ech [e2e-llm-inference-service] o "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n'\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n [e2e-llm-inference-service] 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n [e2e-llm-inference-service] ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common: '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n [e2e-llm-inference-service] '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n [e2e-llm-inference-service] '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n [e2e-llm-inference-service] 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n [e2e-llm-inference-service] 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n ' [e2e-llm-inference-service] read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n [e2e-llm-inference-service] 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-i [e2e-llm-inference-service] nitializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': Non [e2e-llm-inference-service] e},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n [e2e-llm-inference-service] 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n ' [e2e-llm-inference-service] items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_blo [e2e-llm-inference-service] ck_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n [e2e-llm-inference-service] 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n [e2e-llm-inference-service] 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n [e2e-llm-inference-service] 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/v [e2e-llm-inference-service] ar/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '68926',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for lora-single-adapter-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 36, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit' [e2e-llm-inference-service] : {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 54, 36, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 'r [e2e-llm-inference-service] esource_version': '68862',\n 'self_link': None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n [e2e-llm-inference-service] 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n [e2e-llm-inference-service] ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n [e2e-llm-inference-service] 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-acces [e2e-llm-inference-service] s-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n [e2e-llm-inference-service] 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n [e2e-llm-inference-service] 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n [e2e-llm-inference-service] 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_prob [e2e-llm-inference-service] e': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': [e2e-llm-inference-service] 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'g [e2e-llm-inference-service] ce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n [e2e-llm-inference-service] 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n [e2e-llm-inference-service] 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': [e2e-llm-inference-service] datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n [e2e-llm-inference-service] 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n [e2e-llm-inference-service] 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 36, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': [e2e-llm-inference-service] {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 54, 36, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 're [e2e-llm-inference-service] source_version': '68862',\n 'self_link': None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n [e2e-llm-inference-service] 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n [e2e-llm-inference-service] ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n [e2e-llm-inference-service] 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access [e2e-llm-inference-service] -log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n [e2e-llm-inference-service] 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n [e2e-llm-inference-service] 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n [e2e-llm-inference-service] 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe [e2e-llm-inference-service] ': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': [e2e-llm-inference-service] 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gc [e2e-llm-inference-service] e_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n [e2e-llm-inference-service] 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n [e2e-llm-inference-service] 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': d [e2e-llm-inference-service] atetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n [e2e-llm-inference-service] 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n [e2e-llm-inference-service] 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n [e2e-llm-inference-service] 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 36, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n [e2e-llm-inference-service] 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n [e2e-llm-inference-service] 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n [e2e-llm-inference-service] 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {}, [e2e-llm-inference-service] \n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n [e2e-llm-inference-service] 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n [e2e-llm-inference-service] 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n [e2e-llm-inference-service] 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 54, 36, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 'resource_version': '68862',\n 'self_link': None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n [e2e-llm-inference-service] ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'ech [e2e-llm-inference-service] o "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n'\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n [e2e-llm-inference-service] 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n [e2e-llm-inference-service] ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common: '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n [e2e-llm-inference-service] '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n [e2e-llm-inference-service] '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n [e2e-llm-inference-service] 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n [e2e-llm-inference-service] 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n ' [e2e-llm-inference-service] read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n [e2e-llm-inference-service] 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-i [e2e-llm-inference-service] nitializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': Non [e2e-llm-inference-service] e},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n [e2e-llm-inference-service] 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n ' [e2e-llm-inference-service] items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_blo [e2e-llm-inference-service] ck_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n [e2e-llm-inference-service] 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n [e2e-llm-inference-service] 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n [e2e-llm-inference-service] 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/v [e2e-llm-inference-service] ar/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '68967',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for lora-single-adapter-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 36, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit' [e2e-llm-inference-service] : {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 54, 36, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 'r [e2e-llm-inference-service] esource_version': '68862',\n 'self_link': None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n [e2e-llm-inference-service] 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n [e2e-llm-inference-service] ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n [e2e-llm-inference-service] 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-acces [e2e-llm-inference-service] s-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n [e2e-llm-inference-service] 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n [e2e-llm-inference-service] 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n [e2e-llm-inference-service] 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_prob [e2e-llm-inference-service] e': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': [e2e-llm-inference-service] 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'g [e2e-llm-inference-service] ce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n [e2e-llm-inference-service] 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n [e2e-llm-inference-service] 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': [e2e-llm-inference-service] datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n [e2e-llm-inference-service] 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n [e2e-llm-inference-service] 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 36, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': [e2e-llm-inference-service] {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 54, 36, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 're [e2e-llm-inference-service] source_version': '68862',\n 'self_link': None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n [e2e-llm-inference-service] 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n [e2e-llm-inference-service] ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n [e2e-llm-inference-service] 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access [e2e-llm-inference-service] -log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n [e2e-llm-inference-service] 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n [e2e-llm-inference-service] 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n [e2e-llm-inference-service] 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe [e2e-llm-inference-service] ': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': [e2e-llm-inference-service] 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gc [e2e-llm-inference-service] e_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n [e2e-llm-inference-service] 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n [e2e-llm-inference-service] 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': d [e2e-llm-inference-service] atetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n [e2e-llm-inference-service] 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n [e2e-llm-inference-service] 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.61/23"],"mac_address":"0a:58:0a:84:00:3d","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.61/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.61"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:3d",\n'\n ' '\n '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n [e2e-llm-inference-service] 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 36, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-single-adapter-test-kserve-85cd6c88dc-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-single-adapter-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '85cd6c88dc'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"52c7e973-7fc8-4d35-9fdc-3decc731c370"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n [e2e-llm-inference-service] 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n [e2e-llm-inference-service] 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n [e2e-llm-inference-service] 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {}, [e2e-llm-inference-service] \n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n [e2e-llm-inference-service] 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n [e2e-llm-inference-service] 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.61"}': {'.': {},\n 'f:ip': {}}},\n [e2e-llm-inference-service] 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 54, 36, tzinfo=tzlocal())}],\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc-6zqb8',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-single-adapter-test-kserve-85cd6c88dc',\n 'uid': '52c7e973-7fc8-4d35-9fdc-3decc731c370'}],\n 'resource_version': '68862',\n 'self_link': None,\n 'uid': 'd448b029-ed52-4b62-b1cf-efc1e51e2ef3'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n [e2e-llm-inference-service] ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'ech [e2e-llm-inference-service] o "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n'\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n [e2e-llm-inference-service] 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n [e2e-llm-inference-service] ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common: '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n [e2e-llm-inference-service] '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n [e2e-llm-inference-service] '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n [e2e-llm-inference-service] 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n [e2e-llm-inference-service] 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n ' [e2e-llm-inference-service] read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n [e2e-llm-inference-service] 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-i [e2e-llm-inference-service] nitializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': Non [e2e-llm-inference-service] e},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n [e2e-llm-inference-service] 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n ' [e2e-llm-inference-service] items': None,\n 'optional': None,\n 'secret_name': 'lora-single-adapter-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_blo [e2e-llm-inference-service] ck_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-6t5h9',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n [e2e-llm-inference-service] 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 39, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://0437894a66851e7f2b3e8690c8b8e5401da37cf1083c65d04b0fb5e90cd8ae2c',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n [e2e-llm-inference-service] 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n [e2e-llm-inference-service] 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://a14689e0f1ce1baf81270b9e0ebff558c7ba9e1ed225216e2aea3f21c730890b',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 38, 6, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 38, 1, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/v [e2e-llm-inference-service] ar/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-6t5h9',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.61'}],\n 'pod_ip': '10.132.0.61',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 38, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '69004',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [delete_llmisvc] [2026-07-02T13:54:57.033658] end - ✅ in 20.360s [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:189 Cleaning up config router-no-scheduler-lora-single-3cece2ef [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:189 Cleaning up config workload-single-cpu-lora-single-6bc19feb [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:189 Cleaning up config model-fb-opt-125m-with-lora-hf-721eb685 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:44 TIME NAMESPACE SOURCE TYPE REASON MESSAGE [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:45 -------------------------------------------------------------------------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-6599bc8fd-8sr69 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-6599bc8fd-8sr69 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-router-scheduler-6fcbb55f87 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-6599bc8fd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:52 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy auth-disabled-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "auth-disabled-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-disabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-disabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-disabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-disabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:56 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-disabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-67c57d579b-gtwnj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:05 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 27.649s (27.649s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.35:8000/health": dial tcp 10.132.0.35:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-67c57d579b-gtwnj [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.31/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.181s (1.181s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-router-scheduler-688c79755d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-67c57d579b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-enabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-enabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-enabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-enabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-enabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-enabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-6cd6c5884-rrbrr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-6cd6c5884-rrbrr [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-router-scheduler-5b958c76d8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-6cd6c5884 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-invalid-token-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-invalid-token-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-invalid-token-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-invalid-token-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-invalid-token-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:24 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-invalid-token-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:25 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.52:8001/health": dial tcp 10.132.0.52:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.53/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.53:8000/health": dial tcp 10.132.0.53:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-router-scheduler-f7968dfdd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-79d6b9cc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:31 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-pd-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-pd-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:43 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-98989556b-vl89r to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:14:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.45:8000/health": dial tcp 10.132.0.45:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-98989556b-vl89r [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-router-scheduler-579b7545d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-98989556b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:12:57 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:15:01 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/e2e-pvc-model-download-4s522 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:26 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test job-controller Normal SuccessfulCreate Created pod: e2e-pvc-model-download-4s522 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:36 kserve-ci-e2e-test job-controller Normal Completed Job completed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:20 kserve-ci-e2e-test persistentvolume-controller Normal WaitForFirstConsumer waiting for first consumer to be created before binding [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test persistentvolume-controller Normal ExternalProvisioning Waiting for a volume to be created either by the external provisioner 'ebs.csi.aws.com' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal Provisioning External provisioner is provisioning volume for claim "kserve-ci-e2e-test/e2e-pvc-model-storage" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal ProvisioningSucceeded Successfully provisioned volume pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.094s (1.094s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:28 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec0c69dceeb48768325d1a53a749e65786-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.30/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedMount MountVolume.SetUp failed for volume "tls-certs" : secret "gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.60/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:46 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.60:8000/health": dial tcp 10.132.0.60:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.60:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:32 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-3c960099-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-3c960099-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:02 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-3c960099] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" in 808ms (808ms including waiting). Image size: 44914394 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.48:8001/health": dial tcp 10.132.0.48:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.49:8000/health": dial tcp 10.132.0.49:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5b86594bc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:44 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:25:52 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-50bc673d] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:07:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.44:8000/health": dial tcp 10.132.0.44:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:06:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:07:54 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-87882a8e] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 44.193s (44.193s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.62/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:40:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.62:8000/health": dial tcp 10.132.0.62:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 7f6c8cfc6 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:40:17 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test386bd5808c5c450f11fd8632e1f821fb-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:58 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-dc21cb14] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:20 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.45:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:27 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-e95b1dc1] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.42:8000/health": dial tcp 10.132.0.42:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler-67f9b884d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:01 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:19 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-7ca60146] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:06:44 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler-5444b4dbf from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:57 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:06:55 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-ba4d693a] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.55/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.54/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:25 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 686d468674 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:35 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-2577e794-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:27:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:53 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-2577e794] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:56 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.47:8000/health": dial tcp 10.132.0.47:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:51 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-59b9d263-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-59b9d263-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:25 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-59b9d263] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.50:8001/health": dial tcp 10.132.0.50:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.51:8000/health": dial tcp 10.132.0.51:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-79497db4cc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:56 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-e8706282-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:09 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-e8706282] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:30 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.34:8000/health": dial tcp 10.133.0.34:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:36 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:14 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4f8c0978] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.32/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:50 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:48 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-a50492e9] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:40 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-4b931143-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-4b931143-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:21 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-4b931143] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:25 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.42:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-5b1e8f15-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-5b1e8f15-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:22 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-5b1e8f15] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-e45d1f79-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-e45d1f79-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:20 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-e45d1f79] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler-6fcb489785 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.566s (1.566s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-7dbcb75dbc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0, llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc44d181485fad85e662eb092f3749502f-kserve-router-scheduler-69bf65579d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler-5dd88bfbb7 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler-5f555d4d85 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler-54ccf64f4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler-6d86bd4d9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-67b4bb9646 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9, llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler-68cc9685d6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler-749449dbc8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.61/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:39:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-single-adapter-test-kserve-85cd6c88dc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-single-adapter-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-single-adapter-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-single-adapter-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:39:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-single-adapter-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.29/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.483s (1.483s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.29:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 3.651s (3.651s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.88s (1.88s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 2.962s (2.962s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.723s (1.723s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" in 30.447s (30.447s including waiting). Image size: 2989890188 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Liveness probe failed: timeout: failed to connect service "10.132.0.34:9003" within 1s: context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-router-scheduler-77d88bdcf4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-7c94f9d7c6 from 0 to 2 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy precise-prefix-cache-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "precise-prefix-cache-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/precise-prefix-cache-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/precise-prefix-cache-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/precise-prefix-cache-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/precise-prefix-cache-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:15 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [precise-prefix-cache-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:16 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-1-openshift-default-799f46c59b-fnbwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.28/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:31 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "http://10.133.0.28:15021/healthz/ready": context deadline exceeded (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning BackOff Back-off restarting failed container istio-proxy in pod router-gateway-1-openshift-default-799f46c59b-fnbwc_kserve-ci-e2e-test(f244ca92-1cf1-429e-886d-e0d585d0175f) [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:03:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.28:15021/healthz/ready": dial tcp 10.133.0.28:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-1-openshift-default-799f46c59b-fnbwc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:52:58 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-1-openshift-default-799f46c59b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:37 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-2-openshift-default-54c789bdc6-w57g9 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:19 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.44:15021/healthz/ready": dial tcp 10.133.0.44:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-2-openshift-default-54c789bdc6-w57g9 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-2-openshift-default-54c789bdc6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:17 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-2 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.56:8001/health": dial tcp 10.132.0.56:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.57/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.57:8000/health": dial tcp 10.132.0.57:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-prefill-778968b98b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-router-scheduler-5c684fd54d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-694c6ddf4c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:33 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:47 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-587fcd8566-2pqvl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:20:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:29 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-587fcd8566-2pqvl [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-router-scheduler-74545d8489 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-587fcd8566 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:05 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-7c6c58cc96-l62qh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-7c6c58cc96-l62qh [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-router-scheduler-d9ddf6c4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-7c6c58cc96 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy scheduler-inline-config-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "scheduler-inline-config-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/scheduler-inline-config-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/scheduler-inline-config-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/scheduler-inline-config-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/scheduler-inline-config-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [scheduler-inline-config-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-6hzlt to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.59/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:51 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-f2dmd to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.58/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-f2dmd [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-6hzlt [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:07 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy stop-feature-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "stop-feature-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-stop-feature-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/stop-feature-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/stop-feature-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/stop-feature-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [stop-feature-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.InferencePool kserve-ci-e2e-test/stop-feature-test-inference-pool [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:56 kserve-ci-e2e-test LLMInferenceServiceController Warning LLMInferenceServiceNotReady LLMInferenceService [stop-feature-test] is no longer Ready because of: GatewaysReady, HTTPRoutesReady, InferencePoolReady, MainWorkloadReady, PrefillWorkerWorkloadReady, PrefillWorkloadReady, RouterReady, SchedulerWorkloadReady, WorkerWorkloadReady, WorkloadsReady [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_with_lora_adapters] [2026-07-02T13:54:57.761899] end - ❌ 1022.212s: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/lora-single-adapter-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] _____________ test_llm_with_lora_adapters[multiple-lora-adapters] ______________ [e2e-llm-inference-service] [gw1] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), chunked = False [e2e-llm-inference-service] response_conn = [e2e-llm-inference-service] preload_content = False, decode_content = False, enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] # conn.request() calls http.client.*.request, not the method in [e2e-llm-inference-service] # urllib3.request. It also calls makefile (recv) on the socket. [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn.request( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] enforce_content_length=enforce_content_length, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # We are swallowing BrokenPipeError (errno.EPIPE) since the server is [e2e-llm-inference-service] # legitimately able to close the connection after sending a valid response. [e2e-llm-inference-service] # With this behaviour, the received response is still readable. [e2e-llm-inference-service] except BrokenPipeError: [e2e-llm-inference-service] pass [e2e-llm-inference-service] except OSError as e: [e2e-llm-inference-service] # MacOS/Linux [e2e-llm-inference-service] # EPROTOTYPE and ECONNRESET are needed on macOS [e2e-llm-inference-service] # https://erickt.github.io/blog/2014/11/19/adventures-in-debugging-a-potential-osx-kernel-bug/ [e2e-llm-inference-service] # Condition changed later to emit ECONNRESET instead of only EPROTOTYPE. [e2e-llm-inference-service] if e.errno != errno.EPROTOTYPE and e.errno != errno.ECONNRESET: [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset the timeout for the recv() on the socket [e2e-llm-inference-service] read_timeout = timeout_obj.read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn.is_closed: [e2e-llm-inference-service] # In Python 3 socket.py will catch EAGAIN and return None when you [e2e-llm-inference-service] # try and read into the file pointer created by http.client, which [e2e-llm-inference-service] # instead raises a BadStatusLine exception. Instead of catching [e2e-llm-inference-service] # the exception and assuming all BadStatusLine exceptions are read [e2e-llm-inference-service] # timeouts, check for a zero timeout before making the request. [e2e-llm-inference-service] if read_timeout == 0: [e2e-llm-inference-service] raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={read_timeout})" [e2e-llm-inference-service] ) [e2e-llm-inference-service] conn.timeout = read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] # Receive the response from the server [e2e-llm-inference-service] try: [e2e-llm-inference-service] > response = conn.getresponse() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:534: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def getresponse( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] ) -> HTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get the response from the server. [e2e-llm-inference-service] [e2e-llm-inference-service] If the HTTPConnection is in the correct state, returns an instance of HTTPResponse or of whatever object is returned by the response_class variable. [e2e-llm-inference-service] [e2e-llm-inference-service] If a request has not been sent or if a previous response has not be handled, ResponseNotReady is raised. If the HTTP response indicates that the connection should be closed, then it will be closed before the response is returned. When the connection is closed, the underlying socket is closed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Raise the same error as http.client.HTTPConnection [e2e-llm-inference-service] if self._response_options is None: [e2e-llm-inference-service] raise ResponseNotReady() [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset this attribute for being used again. [e2e-llm-inference-service] resp_options = self._response_options [e2e-llm-inference-service] self._response_options = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Since the connection's timeout value may have been updated [e2e-llm-inference-service] # we need to set the timeout on the socket. [e2e-llm-inference-service] self.sock.settimeout(self.timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] # This is needed here to avoid circular import errors [e2e-llm-inference-service] from .response import HTTPResponse [e2e-llm-inference-service] [e2e-llm-inference-service] # Save a reference to the shutdown function before ownership is passed [e2e-llm-inference-service] # to httplib_response [e2e-llm-inference-service] # TODO should we implement it everywhere? [e2e-llm-inference-service] _shutdown = getattr(self.sock, "shutdown", None) [e2e-llm-inference-service] [e2e-llm-inference-service] # Get the response from http.client.HTTPConnection [e2e-llm-inference-service] > httplib_response = super().getresponse() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:571: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def getresponse(self): [e2e-llm-inference-service] """Get the response from the server. [e2e-llm-inference-service] [e2e-llm-inference-service] If the HTTPConnection is in the correct state, returns an [e2e-llm-inference-service] instance of HTTPResponse or of whatever object is returned by [e2e-llm-inference-service] the response_class variable. [e2e-llm-inference-service] [e2e-llm-inference-service] If a request has not been sent or if a previous response has [e2e-llm-inference-service] not be handled, ResponseNotReady is raised. If the HTTP [e2e-llm-inference-service] response indicates that the connection should be closed, then [e2e-llm-inference-service] it will be closed before the response is returned. When the [e2e-llm-inference-service] connection is closed, the underlying socket is closed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] # if a prior response has been completed, then forget about it. [e2e-llm-inference-service] if self.__response and self.__response.isclosed(): [e2e-llm-inference-service] self.__response = None [e2e-llm-inference-service] [e2e-llm-inference-service] # if a prior response exists, then it must be completed (otherwise, we [e2e-llm-inference-service] # cannot read this response's header to determine the connection-close [e2e-llm-inference-service] # behavior) [e2e-llm-inference-service] # [e2e-llm-inference-service] # note: if a prior response existed, but was connection-close, then the [e2e-llm-inference-service] # socket and response were made independent of this HTTPConnection [e2e-llm-inference-service] # object since a new request requires that we open a whole new [e2e-llm-inference-service] # connection [e2e-llm-inference-service] # [e2e-llm-inference-service] # this means the prior response had one of two states: [e2e-llm-inference-service] # 1) will_close: this connection was reset and the prior socket and [e2e-llm-inference-service] # response operate independently [e2e-llm-inference-service] # 2) persistent: the response was retained and we await its [e2e-llm-inference-service] # isclosed() status to become true. [e2e-llm-inference-service] # [e2e-llm-inference-service] if self.__state != _CS_REQ_SENT or self.__response: [e2e-llm-inference-service] raise ResponseNotReady(self.__state) [e2e-llm-inference-service] [e2e-llm-inference-service] if self.debuglevel > 0: [e2e-llm-inference-service] response = self.response_class(self.sock, self.debuglevel, [e2e-llm-inference-service] method=self._method) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = self.response_class(self.sock, method=self._method) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > response.begin() [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:1395: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def begin(self): [e2e-llm-inference-service] if self.headers is not None: [e2e-llm-inference-service] # we've already started reading the response [e2e-llm-inference-service] return [e2e-llm-inference-service] [e2e-llm-inference-service] # read until we get a non-100 response [e2e-llm-inference-service] while True: [e2e-llm-inference-service] > version, status, reason = self._read_status() [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:325: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def _read_status(self): [e2e-llm-inference-service] > line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1") [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:286: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] b = [e2e-llm-inference-service] [e2e-llm-inference-service] def readinto(self, b): [e2e-llm-inference-service] """Read up to len(b) bytes into the writable buffer *b* and return [e2e-llm-inference-service] the number of bytes read. If the socket is non-blocking and no bytes [e2e-llm-inference-service] are available, None is returned. [e2e-llm-inference-service] [e2e-llm-inference-service] If *b* is non-empty, a 0 return value indicates that the connection [e2e-llm-inference-service] was shutdown at the other end. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self._checkClosed() [e2e-llm-inference-service] self._checkReadable() [e2e-llm-inference-service] if self._timeout_occurred: [e2e-llm-inference-service] raise OSError("cannot read from timed out object") [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return self._sock.recv_into(b) [e2e-llm-inference-service] E TimeoutError: timed out [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/socket.py:718: TimeoutError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] > response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:787: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), chunked = False [e2e-llm-inference-service] response_conn = [e2e-llm-inference-service] preload_content = False, decode_content = False, enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] # conn.request() calls http.client.*.request, not the method in [e2e-llm-inference-service] # urllib3.request. It also calls makefile (recv) on the socket. [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn.request( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] enforce_content_length=enforce_content_length, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # We are swallowing BrokenPipeError (errno.EPIPE) since the server is [e2e-llm-inference-service] # legitimately able to close the connection after sending a valid response. [e2e-llm-inference-service] # With this behaviour, the received response is still readable. [e2e-llm-inference-service] except BrokenPipeError: [e2e-llm-inference-service] pass [e2e-llm-inference-service] except OSError as e: [e2e-llm-inference-service] # MacOS/Linux [e2e-llm-inference-service] # EPROTOTYPE and ECONNRESET are needed on macOS [e2e-llm-inference-service] # https://erickt.github.io/blog/2014/11/19/adventures-in-debugging-a-potential-osx-kernel-bug/ [e2e-llm-inference-service] # Condition changed later to emit ECONNRESET instead of only EPROTOTYPE. [e2e-llm-inference-service] if e.errno != errno.EPROTOTYPE and e.errno != errno.ECONNRESET: [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset the timeout for the recv() on the socket [e2e-llm-inference-service] read_timeout = timeout_obj.read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn.is_closed: [e2e-llm-inference-service] # In Python 3 socket.py will catch EAGAIN and return None when you [e2e-llm-inference-service] # try and read into the file pointer created by http.client, which [e2e-llm-inference-service] # instead raises a BadStatusLine exception. Instead of catching [e2e-llm-inference-service] # the exception and assuming all BadStatusLine exceptions are read [e2e-llm-inference-service] # timeouts, check for a zero timeout before making the request. [e2e-llm-inference-service] if read_timeout == 0: [e2e-llm-inference-service] raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={read_timeout})" [e2e-llm-inference-service] ) [e2e-llm-inference-service] conn.timeout = read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] # Receive the response from the server [e2e-llm-inference-service] try: [e2e-llm-inference-service] response = conn.getresponse() [e2e-llm-inference-service] except (BaseSSLError, OSError) as e: [e2e-llm-inference-service] > self._raise_timeout(err=e, url=url, timeout_value=read_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:536: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] err = TimeoutError('timed out') [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] timeout_value = 60 [e2e-llm-inference-service] [e2e-llm-inference-service] def _raise_timeout( [e2e-llm-inference-service] self, [e2e-llm-inference-service] err: BaseSSLError | OSError | SocketTimeout, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] timeout_value: _TYPE_TIMEOUT | None, [e2e-llm-inference-service] ) -> None: [e2e-llm-inference-service] """Is the error actually a timeout? Will raise a ReadTimeout or pass""" [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(err, SocketTimeout): [e2e-llm-inference-service] > raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={timeout_value})" [e2e-llm-inference-service] ) from err [e2e-llm-inference-service] E urllib3.exceptions.ReadTimeoutError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:367: ReadTimeoutError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = , stream = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), verify = '/tmp/ca.crt' [e2e-llm-inference-service] cert = None, proxies = OrderedDict() [e2e-llm-inference-service] [e2e-llm-inference-service] def send( [e2e-llm-inference-service] self, request, stream=False, timeout=None, verify=True, cert=None, proxies=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Sends PreparedRequest object. Returns Response object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param request: The :class:`PreparedRequest ` being sent. [e2e-llm-inference-service] :param stream: (optional) Whether to stream the request content. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple or urllib3 Timeout object [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether [e2e-llm-inference-service] we verify the server's TLS certificate, or a string, in which case it [e2e-llm-inference-service] must be a path to a CA bundle to use [e2e-llm-inference-service] :param cert: (optional) Any user-provided SSL certificate to be trusted. [e2e-llm-inference-service] :param proxies: (optional) The proxies dictionary to apply to the request. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn = self.get_connection_with_tls_context( [e2e-llm-inference-service] request, verify, proxies=proxies, cert=cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] except LocationValueError as e: [e2e-llm-inference-service] raise InvalidURL(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] self.cert_verify(conn, request.url, verify, cert) [e2e-llm-inference-service] url = self.request_url(request, proxies) [e2e-llm-inference-service] self.add_headers( [e2e-llm-inference-service] request, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] verify=verify, [e2e-llm-inference-service] cert=cert, [e2e-llm-inference-service] proxies=proxies, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] chunked = not (request.body is None or "Content-Length" in request.headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(timeout, tuple): [e2e-llm-inference-service] try: [e2e-llm-inference-service] connect, read = timeout [e2e-llm-inference-service] timeout = TimeoutSauce(connect=connect, read=read) [e2e-llm-inference-service] except ValueError: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, " [e2e-llm-inference-service] f"or a single float to set both timeouts to the same value." [e2e-llm-inference-service] ) [e2e-llm-inference-service] elif isinstance(timeout, TimeoutSauce): [e2e-llm-inference-service] pass [e2e-llm-inference-service] else: [e2e-llm-inference-service] timeout = TimeoutSauce(connect=timeout, read=timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > resp = conn.urlopen( [e2e-llm-inference-service] method=request.method, [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] body=request.body, [e2e-llm-inference-service] headers=request.headers, [e2e-llm-inference-service] redirect=False, [e2e-llm-inference-service] assert_same_host=False, [e2e-llm-inference-service] preload_content=False, [e2e-llm-inference-service] decode_content=False, [e2e-llm-inference-service] retries=self.max_retries, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/adapters.py:667: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=7, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=6, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = RemoteDisconnected('Remote end closed connection without response') [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=5, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=4, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=3, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=2, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=1, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "Explain machine learning in simple terms.", "max_tokens": 100}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Length': '104', 'Content-Type': 'application/json'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] > retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:841: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] response = None [e2e-llm-inference-service] error = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] _pool = [e2e-llm-inference-service] _stacktrace = [e2e-llm-inference-service] [e2e-llm-inference-service] def increment( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str | None = None, [e2e-llm-inference-service] url: str | None = None, [e2e-llm-inference-service] response: BaseHTTPResponse | None = None, [e2e-llm-inference-service] error: Exception | None = None, [e2e-llm-inference-service] _pool: ConnectionPool | None = None, [e2e-llm-inference-service] _stacktrace: TracebackType | None = None, [e2e-llm-inference-service] ) -> Self: [e2e-llm-inference-service] """Return a new Retry object with incremented retry counters. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response: A response object, or None, if the server did not [e2e-llm-inference-service] return a response. [e2e-llm-inference-service] :type response: :class:`~urllib3.response.BaseHTTPResponse` [e2e-llm-inference-service] :param Exception error: An error encountered during the request, or [e2e-llm-inference-service] None if the response was received successfully. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: A new ``Retry`` object. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if self.total is False and error: [e2e-llm-inference-service] # Disabled, indicate to re-raise the error. [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] [e2e-llm-inference-service] total = self.total [e2e-llm-inference-service] if total is not None: [e2e-llm-inference-service] total -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] connect = self.connect [e2e-llm-inference-service] read = self.read [e2e-llm-inference-service] redirect = self.redirect [e2e-llm-inference-service] status_count = self.status [e2e-llm-inference-service] other = self.other [e2e-llm-inference-service] cause = "unknown" [e2e-llm-inference-service] status = None [e2e-llm-inference-service] redirect_location = None [e2e-llm-inference-service] [e2e-llm-inference-service] if error and self._is_connection_error(error): [e2e-llm-inference-service] # Connect retry? [e2e-llm-inference-service] if connect is False: [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif connect is not None: [e2e-llm-inference-service] connect -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error and self._is_read_error(error): [e2e-llm-inference-service] # Read retry? [e2e-llm-inference-service] if read is False or method is None or not self._is_method_retryable(method): [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif read is not None: [e2e-llm-inference-service] read -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error: [e2e-llm-inference-service] # Other retry? [e2e-llm-inference-service] if other is not None: [e2e-llm-inference-service] other -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif response and response.get_redirect_location(): [e2e-llm-inference-service] # Redirect retry? [e2e-llm-inference-service] if redirect is not None: [e2e-llm-inference-service] redirect -= 1 [e2e-llm-inference-service] cause = "too many redirects" [e2e-llm-inference-service] response_redirect_location = response.get_redirect_location() [e2e-llm-inference-service] if response_redirect_location: [e2e-llm-inference-service] redirect_location = response_redirect_location [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Incrementing because of a server error like a 500 in [e2e-llm-inference-service] # status_forcelist and the given method is in the allowed_methods [e2e-llm-inference-service] cause = ResponseError.GENERIC_ERROR [e2e-llm-inference-service] if response and response.status: [e2e-llm-inference-service] if status_count is not None: [e2e-llm-inference-service] status_count -= 1 [e2e-llm-inference-service] cause = ResponseError.SPECIFIC_ERROR.format(status_code=response.status) [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] history = self.history + ( [e2e-llm-inference-service] RequestHistory(method, url, error, status, redirect_location), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] new_retry = self.new( [e2e-llm-inference-service] total=total, [e2e-llm-inference-service] connect=connect, [e2e-llm-inference-service] read=read, [e2e-llm-inference-service] redirect=redirect, [e2e-llm-inference-service] status=status_count, [e2e-llm-inference-service] other=other, [e2e-llm-inference-service] history=history, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if new_retry.is_exhausted(): [e2e-llm-inference-service] reason = error or ResponseError(cause) [e2e-llm-inference-service] > raise MaxRetryError(_pool, url, reason) from reason # type: ignore[arg-type] [e2e-llm-inference-service] E urllib3.exceptions.MaxRetryError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/util/retry.py:519: MaxRetryError [e2e-llm-inference-service] [e2e-llm-inference-service] During handling of the above exception, another exception occurred: [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = LoRATestCase(base_refs=['router-no-scheduler', 'workload-single-cpu', 'model-fb-opt-125m-with-multiple-lora'], prompt=..._name='lora-multiple-adapters-test', endpoint='/v1/completions', max_tokens=100, wait_timeout=900, response_timeout=60) [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] LoRATestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-no-scheduler", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="What is Kubernetes?", [e2e-llm-inference-service] expected_adapter_names=["lora-adapter-1"], [e2e-llm-inference-service] service_name="lora-single-adapter-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.llminferenceservice, [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] id="single-lora-adapter-hf", [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] LoRATestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-no-scheduler", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-multiple-lora", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="Explain machine learning in simple terms.", [e2e-llm-inference-service] expected_adapter_names=["lora-adapter-1", "lora-adapter-2"], [e2e-llm-inference-service] service_name="lora-multiple-adapters-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.llminferenceservice, [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] id="multiple-lora-adapters", [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_with_lora_adapters(test_case: LoRATestCase): [e2e-llm-inference-service] """Test LLMInferenceService with LoRA adapters.""" [e2e-llm-inference-service] > run_lora_test(test_case) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_lora_adapters.py:248: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = LoRATestCase(base_refs=['router-no-scheduler', 'workload-single-cpu', 'model-fb-opt-125m-with-multiple-lora'], prompt=..._name='lora-multiple-adapters-test', endpoint='/v1/completions', max_tokens=100, wait_timeout=900, response_timeout=60) [e2e-llm-inference-service] [e2e-llm-inference-service] def run_lora_test(test_case: LoRATestCase): [e2e-llm-inference-service] """Execute a LoRA adapter test case.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] test_id = generate_test_id(test_case) [e2e-llm-inference-service] service_name = test_case.service_name or f"lora-test-{test_id}" [e2e-llm-inference-service] model_name = _get_model_name_from_configs(test_case.base_refs) [e2e-llm-inference-service] created_configs = [] [e2e-llm-inference-service] llm_service = None [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Create unique LLMInferenceServiceConfig resources for each base ref [e2e-llm-inference-service] unique_base_refs = [] [e2e-llm-inference-service] for base_ref in test_case.base_refs: [e2e-llm-inference-service] if base_ref not in LLMINFERENCESERVICE_CONFIGS: [e2e-llm-inference-service] raise ValueError(f"Unknown base reference: {base_ref}") [e2e-llm-inference-service] [e2e-llm-inference-service] # Generate unique config name to avoid conflicts in parallel test runs [e2e-llm-inference-service] unique_config_name = generate_k8s_safe_suffix(base_ref, [service_name]) [e2e-llm-inference-service] unique_base_refs.append(unique_config_name) [e2e-llm-inference-service] [e2e-llm-inference-service] config = LLMINFERENCESERVICE_CONFIGS[base_ref] [e2e-llm-inference-service] config_body = { [e2e-llm-inference-service] "apiVersion": "serving.kserve.io/v1alpha1", [e2e-llm-inference-service] "kind": "LLMInferenceServiceConfig", [e2e-llm-inference-service] "metadata": { [e2e-llm-inference-service] "name": unique_config_name, [e2e-llm-inference-service] "namespace": KSERVE_TEST_NAMESPACE, [e2e-llm-inference-service] }, [e2e-llm-inference-service] "spec": config, [e2e-llm-inference-service] } [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info("Creating LLMInferenceServiceConfig: %s", unique_config_name) [e2e-llm-inference-service] _create_or_update_llmisvc_config( [e2e-llm-inference-service] kserve_client, config_body, KSERVE_TEST_NAMESPACE [e2e-llm-inference-service] ) [e2e-llm-inference-service] created_configs.append(unique_config_name) [e2e-llm-inference-service] [e2e-llm-inference-service] # Create the service with unique base refs [e2e-llm-inference-service] llm_service = build_llm_service_from_refs(service_name, unique_base_refs) [e2e-llm-inference-service] if not llm_service.metadata.annotations: [e2e-llm-inference-service] llm_service.metadata.annotations = {} [e2e-llm-inference-service] llm_service.metadata.annotations["security.opendatahub.io/enable-auth"] = ( [e2e-llm-inference-service] "false" [e2e-llm-inference-service] ) [e2e-llm-inference-service] logger.info("Creating LLMInferenceService: %s", service_name) [e2e-llm-inference-service] create_llmisvc(kserve_client, llm_service) [e2e-llm-inference-service] [e2e-llm-inference-service] # Wait for service to be ready [e2e-llm-inference-service] logger.info("Waiting for service %s to be ready...", service_name) [e2e-llm-inference-service] wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, llm_service, timeout_seconds=test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Get inference URL [e2e-llm-inference-service] base_url = get_llm_service_url(kserve_client, llm_service) [e2e-llm-inference-service] inference_url = base_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] # Test base model inference [e2e-llm-inference-service] logger.info("Testing base model inference...") [e2e-llm-inference-service] base_payload = { [e2e-llm-inference-service] "model": model_name, [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] [e2e-llm-inference-service] > base_response = post_with_retry( [e2e-llm-inference-service] inference_url, [e2e-llm-inference-service] json_data=base_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_lora_adapters.py:150: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] [e2e-llm-inference-service] def post_with_retry( [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] *, [e2e-llm-inference-service] headers: Dict = None, [e2e-llm-inference-service] json_data: Union[Dict, List] = None, [e2e-llm-inference-service] data: Union[str, bytes] = None, [e2e-llm-inference-service] stream: bool = False, [e2e-llm-inference-service] timeout: float = None, [e2e-llm-inference-service] total_retries: int = DEFAULT_RETRY_TOTAL, [e2e-llm-inference-service] backoff_factor: float = DEFAULT_RETRY_BACKOFF_FACTOR, [e2e-llm-inference-service] retry_status_codes=DEFAULT_RETRY_STATUS_CODES, [e2e-llm-inference-service] ) -> requests.Response: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Send POST request with retries for transient HTTP and network failures. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if json_data is not None and data is not None: [e2e-llm-inference-service] raise ValueError("Only one of json_data or data can be provided.") [e2e-llm-inference-service] [e2e-llm-inference-service] with _retry_session( [e2e-llm-inference-service] ["POST"], total_retries, backoff_factor, retry_status_codes [e2e-llm-inference-service] ) as session: [e2e-llm-inference-service] > return session.post( [e2e-llm-inference-service] url, [e2e-llm-inference-service] json=json_data, [e2e-llm-inference-service] data=data, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] common/http_retry.py:70: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] data = None [e2e-llm-inference-service] json = {'max_tokens': 100, 'model': 'facebook/opt-125m', 'prompt': 'Explain machine learning in simple terms.'} [e2e-llm-inference-service] kwargs = {'headers': None, 'stream': False, 'timeout': 60} [e2e-llm-inference-service] [e2e-llm-inference-service] def post(self, url, data=None, json=None, **kwargs): [e2e-llm-inference-service] r"""Sends a POST request. Returns :class:`Response` object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: URL for the new :class:`Request` object. [e2e-llm-inference-service] :param data: (optional) Dictionary, list of tuples, bytes, or file-like [e2e-llm-inference-service] object to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param json: (optional) json to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param \*\*kwargs: Optional arguments that ``request`` takes. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.request("POST", url, data=data, json=json, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:637: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = , method = 'POST' [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions' [e2e-llm-inference-service] params = None, data = None, headers = None, cookies = None, files = None [e2e-llm-inference-service] auth = None, timeout = 60, allow_redirects = True, proxies = {}, hooks = None [e2e-llm-inference-service] stream = False, verify = None, cert = None [e2e-llm-inference-service] json = {'max_tokens': 100, 'model': 'facebook/opt-125m', 'prompt': 'Explain machine learning in simple terms.'} [e2e-llm-inference-service] [e2e-llm-inference-service] def request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] params=None, [e2e-llm-inference-service] data=None, [e2e-llm-inference-service] headers=None, [e2e-llm-inference-service] cookies=None, [e2e-llm-inference-service] files=None, [e2e-llm-inference-service] auth=None, [e2e-llm-inference-service] timeout=None, [e2e-llm-inference-service] allow_redirects=True, [e2e-llm-inference-service] proxies=None, [e2e-llm-inference-service] hooks=None, [e2e-llm-inference-service] stream=None, [e2e-llm-inference-service] verify=None, [e2e-llm-inference-service] cert=None, [e2e-llm-inference-service] json=None, [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Constructs a :class:`Request `, prepares it and sends it. [e2e-llm-inference-service] Returns :class:`Response ` object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: method for the new :class:`Request` object. [e2e-llm-inference-service] :param url: URL for the new :class:`Request` object. [e2e-llm-inference-service] :param params: (optional) Dictionary or bytes to be sent in the query [e2e-llm-inference-service] string for the :class:`Request`. [e2e-llm-inference-service] :param data: (optional) Dictionary, list of tuples, bytes, or file-like [e2e-llm-inference-service] object to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param json: (optional) json to send in the body of the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param headers: (optional) Dictionary of HTTP Headers to send with the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param cookies: (optional) Dict or CookieJar object to send with the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param files: (optional) Dictionary of ``'filename': file-like-objects`` [e2e-llm-inference-service] for multipart encoding upload. [e2e-llm-inference-service] :param auth: (optional) Auth tuple or callable to enable [e2e-llm-inference-service] Basic/Digest/Custom HTTP Auth. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple [e2e-llm-inference-service] :param allow_redirects: (optional) Set to True by default. [e2e-llm-inference-service] :type allow_redirects: bool [e2e-llm-inference-service] :param proxies: (optional) Dictionary mapping protocol or protocol and [e2e-llm-inference-service] hostname to the URL of the proxy. [e2e-llm-inference-service] :param hooks: (optional) Dictionary mapping hook name to one event or [e2e-llm-inference-service] list of events, event must be callable. [e2e-llm-inference-service] :param stream: (optional) whether to immediately download the response [e2e-llm-inference-service] content. Defaults to ``False``. [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether we verify [e2e-llm-inference-service] the server's TLS certificate, or a string, in which case it must be a path [e2e-llm-inference-service] to a CA bundle to use. Defaults to ``True``. When set to [e2e-llm-inference-service] ``False``, requests will accept any TLS certificate presented by [e2e-llm-inference-service] the server, and will ignore hostname mismatches and/or expired [e2e-llm-inference-service] certificates, which will make your application vulnerable to [e2e-llm-inference-service] man-in-the-middle (MitM) attacks. Setting verify to ``False`` [e2e-llm-inference-service] may be useful during local development or testing. [e2e-llm-inference-service] :param cert: (optional) if String, path to ssl client cert file (.pem). [e2e-llm-inference-service] If Tuple, ('cert', 'key') pair. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Create the Request. [e2e-llm-inference-service] req = Request( [e2e-llm-inference-service] method=method.upper(), [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] files=files, [e2e-llm-inference-service] data=data or {}, [e2e-llm-inference-service] json=json, [e2e-llm-inference-service] params=params or {}, [e2e-llm-inference-service] auth=auth, [e2e-llm-inference-service] cookies=cookies, [e2e-llm-inference-service] hooks=hooks, [e2e-llm-inference-service] ) [e2e-llm-inference-service] prep = self.prepare_request(req) [e2e-llm-inference-service] [e2e-llm-inference-service] proxies = proxies or {} [e2e-llm-inference-service] [e2e-llm-inference-service] settings = self.merge_environment_settings( [e2e-llm-inference-service] prep.url, proxies, stream, verify, cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Send the request. [e2e-llm-inference-service] send_kwargs = { [e2e-llm-inference-service] "timeout": timeout, [e2e-llm-inference-service] "allow_redirects": allow_redirects, [e2e-llm-inference-service] } [e2e-llm-inference-service] send_kwargs.update(settings) [e2e-llm-inference-service] > resp = self.send(prep, **send_kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:589: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = [e2e-llm-inference-service] kwargs = {'cert': None, 'proxies': OrderedDict(), 'stream': False, 'timeout': 60, ...} [e2e-llm-inference-service] allow_redirects = True, stream = False, hooks = {'response': []} [e2e-llm-inference-service] adapter = [e2e-llm-inference-service] start = 1783000602.5952365 [e2e-llm-inference-service] [e2e-llm-inference-service] def send(self, request, **kwargs): [e2e-llm-inference-service] """Send a given PreparedRequest. [e2e-llm-inference-service] [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Set defaults that the hooks can utilize to ensure they always have [e2e-llm-inference-service] # the correct parameters to reproduce the previous request. [e2e-llm-inference-service] kwargs.setdefault("stream", self.stream) [e2e-llm-inference-service] kwargs.setdefault("verify", self.verify) [e2e-llm-inference-service] kwargs.setdefault("cert", self.cert) [e2e-llm-inference-service] if "proxies" not in kwargs: [e2e-llm-inference-service] kwargs["proxies"] = resolve_proxies(request, self.proxies, self.trust_env) [e2e-llm-inference-service] [e2e-llm-inference-service] # It's possible that users might accidentally send a Request object. [e2e-llm-inference-service] # Guard against that specific failure case. [e2e-llm-inference-service] if isinstance(request, Request): [e2e-llm-inference-service] raise ValueError("You can only send PreparedRequests.") [e2e-llm-inference-service] [e2e-llm-inference-service] # Set up variables needed for resolve_redirects and dispatching of hooks [e2e-llm-inference-service] allow_redirects = kwargs.pop("allow_redirects", True) [e2e-llm-inference-service] stream = kwargs.get("stream") [e2e-llm-inference-service] hooks = request.hooks [e2e-llm-inference-service] [e2e-llm-inference-service] # Get the appropriate adapter to use [e2e-llm-inference-service] adapter = self.get_adapter(url=request.url) [e2e-llm-inference-service] [e2e-llm-inference-service] # Start time (approximately) of the request [e2e-llm-inference-service] start = preferred_clock() [e2e-llm-inference-service] [e2e-llm-inference-service] # Send the request [e2e-llm-inference-service] > r = adapter.send(request, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:703: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = , stream = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), verify = '/tmp/ca.crt' [e2e-llm-inference-service] cert = None, proxies = OrderedDict() [e2e-llm-inference-service] [e2e-llm-inference-service] def send( [e2e-llm-inference-service] self, request, stream=False, timeout=None, verify=True, cert=None, proxies=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Sends PreparedRequest object. Returns Response object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param request: The :class:`PreparedRequest ` being sent. [e2e-llm-inference-service] :param stream: (optional) Whether to stream the request content. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple or urllib3 Timeout object [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether [e2e-llm-inference-service] we verify the server's TLS certificate, or a string, in which case it [e2e-llm-inference-service] must be a path to a CA bundle to use [e2e-llm-inference-service] :param cert: (optional) Any user-provided SSL certificate to be trusted. [e2e-llm-inference-service] :param proxies: (optional) The proxies dictionary to apply to the request. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn = self.get_connection_with_tls_context( [e2e-llm-inference-service] request, verify, proxies=proxies, cert=cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] except LocationValueError as e: [e2e-llm-inference-service] raise InvalidURL(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] self.cert_verify(conn, request.url, verify, cert) [e2e-llm-inference-service] url = self.request_url(request, proxies) [e2e-llm-inference-service] self.add_headers( [e2e-llm-inference-service] request, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] verify=verify, [e2e-llm-inference-service] cert=cert, [e2e-llm-inference-service] proxies=proxies, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] chunked = not (request.body is None or "Content-Length" in request.headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(timeout, tuple): [e2e-llm-inference-service] try: [e2e-llm-inference-service] connect, read = timeout [e2e-llm-inference-service] timeout = TimeoutSauce(connect=connect, read=read) [e2e-llm-inference-service] except ValueError: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, " [e2e-llm-inference-service] f"or a single float to set both timeouts to the same value." [e2e-llm-inference-service] ) [e2e-llm-inference-service] elif isinstance(timeout, TimeoutSauce): [e2e-llm-inference-service] pass [e2e-llm-inference-service] else: [e2e-llm-inference-service] timeout = TimeoutSauce(connect=timeout, read=timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] resp = conn.urlopen( [e2e-llm-inference-service] method=request.method, [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] body=request.body, [e2e-llm-inference-service] headers=request.headers, [e2e-llm-inference-service] redirect=False, [e2e-llm-inference-service] assert_same_host=False, [e2e-llm-inference-service] preload_content=False, [e2e-llm-inference-service] decode_content=False, [e2e-llm-inference-service] retries=self.max_retries, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] except (ProtocolError, OSError) as err: [e2e-llm-inference-service] raise ConnectionError(err, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] except MaxRetryError as e: [e2e-llm-inference-service] if isinstance(e.reason, ConnectTimeoutError): [e2e-llm-inference-service] # TODO: Remove this in 3.0.0: see #2811 [e2e-llm-inference-service] if not isinstance(e.reason, NewConnectionError): [e2e-llm-inference-service] raise ConnectTimeout(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, ResponseError): [e2e-llm-inference-service] raise RetryError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, _ProxyError): [e2e-llm-inference-service] raise ProxyError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, _SSLError): [e2e-llm-inference-service] # This branch is for urllib3 v1.22 and later. [e2e-llm-inference-service] raise SSLError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] > raise ConnectionError(e, request=request) [e2e-llm-inference-service] E requests.exceptions.ConnectionError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/adapters.py:700: ConnectionError [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_with_lora_adapters] [2026-07-02T13:54:58.163122] start - args=(), kwargs={'test_case': LoRATestCase(base_refs=['router-no-scheduler', 'workload-single-cpu', 'model-fb-opt-125m-with-multiple-lora'], prompt='Explain machine learning in simple terms.', expected_adapter_names=['lora-adapter-1', 'lora-adapter-2'], service_name='lora-multiple-adapters-test', endpoint='/v1/completions', max_tokens=100, wait_timeout=900, response_timeout=60)} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:116 Creating LLMInferenceServiceConfig: router-no-scheduler-lora-multip-343dff1f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig router-no-scheduler-lora-multip-343dff1f in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig router-no-scheduler-lora-multip-343dff1f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig router-no-scheduler-lora-multip-343dff1f [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:116 Creating LLMInferenceServiceConfig: workload-single-cpu-lora-multip-b14d2ed8 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig workload-single-cpu-lora-multip-b14d2ed8 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig workload-single-cpu-lora-multip-b14d2ed8 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig workload-single-cpu-lora-multip-b14d2ed8 [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:116 Creating LLMInferenceServiceConfig: model-fb-opt-125m-with-multiple-7cbd2065 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig model-fb-opt-125m-with-multiple-7cbd2065 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig model-fb-opt-125m-with-multiple-7cbd2065 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig model-fb-opt-125m-with-multiple-7cbd2065 [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:129 Creating LLMInferenceService: lora-multiple-adapters-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-02T13:54:58.274185] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'lora-multiple-adapters-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-no-scheduler-lora-multip-343dff1f'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-lora-multip-b14d2ed8'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-with-multiple-7cbd2065'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-02T13:54:58.306399] end - ✅ in 0.032s [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:133 Waiting for service lora-multiple-adapters-test to be ready... [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-02T13:54:58.306729] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'lora-multiple-adapters-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-no-scheduler-lora-multip-343dff1f'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-lora-multip-b14d2ed8'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-with-multiple-7cbd2065'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={'timeout_seconds': 900} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-02T13:55:11Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route: v1.HTTPRouteStatus{RouteStatus:v1.RouteStatus{Parents:[]v1.RouteParentStatus(nil)}}]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:55:11Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:55:11Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:55:11Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route: v1.HTTPRouteStatus{RouteStatus:v1.RouteStatus{Parents:[]v1.RouteParentStatus(nil)}}]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:55:11Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route: v1.HTTPRouteStatus{RouteStatus:v1.RouteStatus{Parents:[]v1.RouteParentStatus(nil)}}]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:55:11Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-02T13:55:17Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:55:17Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:55:11Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:55:17Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:55:17Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:55:17Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [wait_for_llm_isvc_ready] [2026-07-02T13:56:42.584443] end - ✅ in 104.277s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [get_llm_service_url] [2026-07-02T13:56:42.584676] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'lora-multiple-adapters-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-no-scheduler-lora-multip-343dff1f'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-lora-multip-b14d2ed8'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-with-multiple-7cbd2065'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [get_llm_service_url] [2026-07-02T13:56:42.593269] end - ✅ in 0.008s [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:143 Testing base model inference... [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=7, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=6, connect=None, read=None, redirect=None, status=None)) after connection broken by 'RemoteDisconnected('Remote end closed connection without response')': /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=5, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=4, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=3, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=2, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=1, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=0, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:181 Cleaning up service lora-multiple-adapters-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [delete_llmisvc] [2026-07-02T14:11:00.640212] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'lora-multiple-adapters-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-no-scheduler-lora-multip-343dff1f'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-lora-multip-b14d2ed8'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-with-multiple-7cbd2065'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for lora-multiple-adapters-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"n [e2e-llm-inference-service] ame":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': { [e2e-llm-inference-service] },\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n [e2e-llm-inference-service] 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n [e2e-llm-inference-service] 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n 'resource_version': '70400',\n [e2e-llm-inference-service] 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n [e2e-llm-inference-service] '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_g [e2e-llm-inference-service] id_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n [e2e-llm-inference-service] 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n [e2e-llm-inference-service] ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'val [e2e-llm-inference-service] ue': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': [e2e-llm-inference-service] None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n [e2e-llm-inference-service] 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n [e2e-llm-inference-service] 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem' [e2e-llm-inference-service] : None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': N [e2e-llm-inference-service] one,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': N [e2e-llm-inference-service] one,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n [e2e-llm-inference-service] 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs': None,\n 'persistent_v [e2e-llm-inference-service] olume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'sta [e2e-llm-inference-service] tus': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read [e2e-llm-inference-service] _only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n 'waiting': None},\n [e2e-llm-inference-service] 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"na [e2e-llm-inference-service] me":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {} [e2e-llm-inference-service] ,\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n [e2e-llm-inference-service] 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n [e2e-llm-inference-service] 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n 'resource_version': '70400',\n [e2e-llm-inference-service] 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n [e2e-llm-inference-service] '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gi [e2e-llm-inference-service] d_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n ' [e2e-llm-inference-service] tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n [e2e-llm-inference-service] ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'valu [e2e-llm-inference-service] e': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': N [e2e-llm-inference-service] one,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n [e2e-llm-inference-service] 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n [e2e-llm-inference-service] 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': [e2e-llm-inference-service] None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': No [e2e-llm-inference-service] ne,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': No [e2e-llm-inference-service] ne,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n [e2e-llm-inference-service] 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs': None,\n 'persistent_vo [e2e-llm-inference-service] lume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'stat [e2e-llm-inference-service] us': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_ [e2e-llm-inference-service] only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n 'waiting': None},\n [e2e-llm-inference-service] 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n '" [e2e-llm-inference-service] default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n [e2e-llm-inference-service] 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:li [e2e-llm-inference-service] mits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n [e2e-llm-inference-service] 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name' [e2e-llm-inference-service] : {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n [e2e-llm-inference-service] 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fie [e2e-llm-inference-service] lds_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n [e2e-llm-inference-service] 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n 'resource_version': '70400',\n 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n [e2e-llm-inference-service] '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any '\n 'active HCAs\n'\n [e2e-llm-inference-service] ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'echo "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n'\n ' declare -A '\n [e2e-llm-inference-service] 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in [e2e-llm-inference-service] '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common: '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NV [e2e-llm-inference-service] SHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSIO [e2e-llm-inference-service] N" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'v [e2e-llm-inference-service] alue': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n [e2e-llm-inference-service] 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_ [e2e-llm-inference-service] pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n [e2e-llm-inference-service] 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n [e2e-llm-inference-service] 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n [e2e-llm-inference-service] 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n [e2e-llm-inference-service] 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n [e2e-llm-inference-service] 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n [e2e-llm-inference-service] 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluste [e2e-llm-inference-service] r_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'message': None,\n 'reason' [e2e-llm-inference-service] : None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n [e2e-llm-inference-service] 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n [e2e-llm-inference-service] 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '82440',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for lora-multiple-adapters-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:si [e2e-llm-inference-service] zeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n [e2e-llm-inference-service] 'resource_version': '82455',\n 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_ [e2e-llm-inference-service] state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n [e2e-llm-inference-service] ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback [e2e-llm-inference-service] if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n [e2e-llm-inference-service] ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEV [e2e-llm-inference-service] EL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n [e2e-llm-inference-service] 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n [e2e-llm-inference-service] 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n [e2e-llm-inference-service] 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n [e2e-llm-inference-service] 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': [e2e-llm-inference-service] None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n [e2e-llm-inference-service] 'optional': None,\n 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs' [e2e-llm-inference-service] : None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'rea [e2e-llm-inference-service] son': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n [e2e-llm-inference-service] 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:siz [e2e-llm-inference-service] eLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n [e2e-llm-inference-service] 'resource_version': '82455',\n 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_s [e2e-llm-inference-service] tate_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n [e2e-llm-inference-service] ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback [e2e-llm-inference-service] if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n [e2e-llm-inference-service] ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVE [e2e-llm-inference-service] L',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n [e2e-llm-inference-service] 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n [e2e-llm-inference-service] 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n [e2e-llm-inference-service] 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n [e2e-llm-inference-service] 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': N [e2e-llm-inference-service] one,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n ' [e2e-llm-inference-service] optional': None,\n 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs': [e2e-llm-inference-service] None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'reas [e2e-llm-inference-service] on': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n ' [e2e-llm-inference-service] waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n [e2e-llm-inference-service] '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n [e2e-llm-inference-service] 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n [e2e-llm-inference-service] 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n [e2e-llm-inference-service] 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n [e2e-llm-inference-service] 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n [e2e-llm-inference-service] {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n [e2e-llm-inference-service] 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n 'resource_version': '82455',\n 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/t [e2e-llm-inference-service] ypes/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if [e2e-llm-inference-service] we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'echo "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n [e2e-llm-inference-service] '\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n [e2e-llm-inference-service] ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common: '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' [e2e-llm-inference-service] export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${ [e2e-llm-inference-service] VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n [e2e-llm-inference-service] {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': [e2e-llm-inference-service] None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n [e2e-llm-inference-service] 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n [e2e-llm-inference-service] {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecut [e2e-llm-inference-service] e',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None, [e2e-llm-inference-service] \n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'do [e2e-llm-inference-service] wnward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n [e2e-llm-inference-service] 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_ac [e2e-llm-inference-service] count_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n [e2e-llm-inference-service] 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n [e2e-llm-inference-service] 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n 'waiting': None},\n [e2e-llm-inference-service] 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '82508',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for lora-multiple-adapters-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:si [e2e-llm-inference-service] zeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n [e2e-llm-inference-service] 'resource_version': '82455',\n 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_ [e2e-llm-inference-service] state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n [e2e-llm-inference-service] ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback [e2e-llm-inference-service] if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n [e2e-llm-inference-service] ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEV [e2e-llm-inference-service] EL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n [e2e-llm-inference-service] 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n [e2e-llm-inference-service] 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n [e2e-llm-inference-service] 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n [e2e-llm-inference-service] 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': [e2e-llm-inference-service] None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n [e2e-llm-inference-service] 'optional': None,\n 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs' [e2e-llm-inference-service] : None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'rea [e2e-llm-inference-service] son': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n [e2e-llm-inference-service] 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:siz [e2e-llm-inference-service] eLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n [e2e-llm-inference-service] 'resource_version': '82455',\n 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_s [e2e-llm-inference-service] tate_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n [e2e-llm-inference-service] ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback [e2e-llm-inference-service] if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n [e2e-llm-inference-service] ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVE [e2e-llm-inference-service] L',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n [e2e-llm-inference-service] 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n [e2e-llm-inference-service] 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n [e2e-llm-inference-service] 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n [e2e-llm-inference-service] 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': N [e2e-llm-inference-service] one,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n ' [e2e-llm-inference-service] optional': None,\n 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs': [e2e-llm-inference-service] None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'reas [e2e-llm-inference-service] on': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n ' [e2e-llm-inference-service] waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n [e2e-llm-inference-service] '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n [e2e-llm-inference-service] 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n [e2e-llm-inference-service] 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n [e2e-llm-inference-service] 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n [e2e-llm-inference-service] 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n [e2e-llm-inference-service] {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n [e2e-llm-inference-service] 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n 'resource_version': '82455',\n 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/t [e2e-llm-inference-service] ypes/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if [e2e-llm-inference-service] we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'echo "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n [e2e-llm-inference-service] '\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n [e2e-llm-inference-service] ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common: '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' [e2e-llm-inference-service] export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${ [e2e-llm-inference-service] VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n [e2e-llm-inference-service] {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': [e2e-llm-inference-service] None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n [e2e-llm-inference-service] 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n [e2e-llm-inference-service] {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecut [e2e-llm-inference-service] e',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None, [e2e-llm-inference-service] \n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'do [e2e-llm-inference-service] wnward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n [e2e-llm-inference-service] 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_ac [e2e-llm-inference-service] count_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n [e2e-llm-inference-service] 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n [e2e-llm-inference-service] 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n 'waiting': None},\n [e2e-llm-inference-service] 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '82544',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for lora-multiple-adapters-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:si [e2e-llm-inference-service] zeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n [e2e-llm-inference-service] 'resource_version': '82455',\n 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_ [e2e-llm-inference-service] state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n [e2e-llm-inference-service] ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback [e2e-llm-inference-service] if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n [e2e-llm-inference-service] ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEV [e2e-llm-inference-service] EL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n [e2e-llm-inference-service] 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n [e2e-llm-inference-service] 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n [e2e-llm-inference-service] 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n [e2e-llm-inference-service] 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': [e2e-llm-inference-service] None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n [e2e-llm-inference-service] 'optional': None,\n 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs' [e2e-llm-inference-service] : None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'rea [e2e-llm-inference-service] son': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n [e2e-llm-inference-service] 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n [e2e-llm-inference-service] 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n [e2e-llm-inference-service] 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n [e2e-llm-inference-service] 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:siz [e2e-llm-inference-service] eLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n [e2e-llm-inference-service] 'resource_version': '82455',\n 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_s [e2e-llm-inference-service] tate_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n [e2e-llm-inference-service] ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback [e2e-llm-inference-service] if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n [e2e-llm-inference-service] ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVE [e2e-llm-inference-service] L',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n [e2e-llm-inference-service] 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n [e2e-llm-inference-service] 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n [e2e-llm-inference-service] 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n [e2e-llm-inference-service] 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': N [e2e-llm-inference-service] one,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n ' [e2e-llm-inference-service] optional': None,\n 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs': [e2e-llm-inference-service] None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'reas [e2e-llm-inference-service] on': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n ' [e2e-llm-inference-service] waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.63/23"],"mac_address":"0a:58:0a:84:00:3f","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.63/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.63"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:3f",\n'\n ' '\n [e2e-llm-inference-service] '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'lora-multiple-adapters-test-kserve-ccbb969bd-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'lora-multiple-adapters-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': 'ccbb969bd'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"6aee7170-295a-4980-bb4f-fd0a01634460"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n [e2e-llm-inference-service] 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n [e2e-llm-inference-service] 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n [e2e-llm-inference-service] 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"CA_BUNDLE_CONFIGMAP_NAME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"CA_BUNDLE_VOLUME_MOUNT_POINT"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/etc/ssl/custom-certs"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/mnt"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"cabundle-cert"}': {'.': {},\n 'f:configMap': {'.': {},\n 'f:defaultMode': {},\n [e2e-llm-inference-service] 'f:name': {}},\n 'f:name': {}},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n [e2e-llm-inference-service] {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n [e2e-llm-inference-service] 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.63"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, tzinfo=tzlocal())}],\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd-65ln7',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'lora-multiple-adapters-test-kserve-ccbb969bd',\n 'uid': '6aee7170-295a-4980-bb4f-fd0a01634460'}],\n 'resource_version': '82455',\n 'self_link': None,\n 'uid': '1afe24a3-7144-489a-8d6e-d02a7f1678ef'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--enable-lora',\n '--max-lora-rank=64',\n '--max-loras=2',\n '--max-cpu-loras=4',\n '--lora-modules',\n '\'{"name":"lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-1","path":"/mnt/lora/lora-adapter-1"}\'',\n '\'{"name":"lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\'',\n '\'{"name":"publishers/kserve-ci-e2e-test/models/lora-adapter-2","path":"/mnt/lora/lora-adapter-2"}\''],\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/t [e2e-llm-inference-service] ypes/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if [e2e-llm-inference-service] we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'echo "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n [e2e-llm-inference-service] '\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n [e2e-llm-inference-service] ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common: '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' [e2e-llm-inference-service] export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${ [e2e-llm-inference-service] VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n [e2e-llm-inference-service] {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': [e2e-llm-inference-service] None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n [e2e-llm-inference-service] 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n [e2e-llm-inference-service] 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-1',\n 'hf://edbeeching/opt-125m-lora',\n '/mnt/lora/lora-adapter-2'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n [e2e-llm-inference-service] {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'CA_BUNDLE_CONFIGMAP_NAME',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'CA_BUNDLE_VOLUME_MOUNT_POINT',\n 'value': '/etc/ssl/custom-certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'mount_propagation': None,\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecut [e2e-llm-inference-service] e',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None, [e2e-llm-inference-service] \n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'do [e2e-llm-inference-service] wnward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'lora-multiple-adapters-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': {'default_mode': 420,\n 'items': None,\n 'name': 'odh-kserve-custom-ca-bundle',\n 'optional': None},\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n [e2e-llm-inference-service] 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'cabundle-cert',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-7x75z',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_ac [e2e-llm-inference-service] count_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 56, 41, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal()),\n [e2e-llm-inference-service] 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://e8672e6e5855639eaa3b04e615fe2971527a5284b89fd554a1a715790ab46f12',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 13, 55, 18, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n [e2e-llm-inference-service] 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d2b42f8d2e5a6f40fcfcde6284201fada107c96b40130ebe767a96d0d873c8a8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 13, 55, 17, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 13, 55, 12, tzinfo=tzlocal())},\n 'waiting': None},\n [e2e-llm-inference-service] 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/etc/ssl/custom-certs',\n 'name': 'cabundle-cert',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-7x75z',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.63'}],\n 'pod_ip': '10.132.0.63',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 13, 55, 11, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '82590',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [delete_llmisvc] [2026-07-02T14:11:21.012190] end - ✅ in 20.371s [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:189 Cleaning up config router-no-scheduler-lora-multip-343dff1f [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:189 Cleaning up config workload-single-cpu-lora-multip-b14d2ed8 [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_lora_adapters:test_llm_lora_adapters.py:189 Cleaning up config model-fb-opt-125m-with-multiple-7cbd2065 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:44 TIME NAMESPACE SOURCE TYPE REASON MESSAGE [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:45 -------------------------------------------------------------------------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-6599bc8fd-8sr69 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-6599bc8fd-8sr69 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-router-scheduler-6fcbb55f87 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-6599bc8fd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:52 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy auth-disabled-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "auth-disabled-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-disabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-disabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-disabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-disabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:56 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-disabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-67c57d579b-gtwnj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:05 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 27.649s (27.649s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.35:8000/health": dial tcp 10.132.0.35:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-67c57d579b-gtwnj [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.31/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.181s (1.181s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-router-scheduler-688c79755d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-67c57d579b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-enabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-enabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-enabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-enabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-enabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-enabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-6cd6c5884-rrbrr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-6cd6c5884-rrbrr [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-router-scheduler-5b958c76d8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-6cd6c5884 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-invalid-token-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-invalid-token-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-invalid-token-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-invalid-token-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-invalid-token-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:24 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-invalid-token-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:25 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.52:8001/health": dial tcp 10.132.0.52:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.53/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.53:8000/health": dial tcp 10.132.0.53:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-router-scheduler-f7968dfdd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-79d6b9cc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:31 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-pd-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-pd-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:43 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-98989556b-vl89r to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:14:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.45:8000/health": dial tcp 10.132.0.45:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-98989556b-vl89r [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-router-scheduler-579b7545d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-98989556b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:12:57 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:15:01 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/e2e-pvc-model-download-4s522 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:26 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test job-controller Normal SuccessfulCreate Created pod: e2e-pvc-model-download-4s522 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:36 kserve-ci-e2e-test job-controller Normal Completed Job completed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:20 kserve-ci-e2e-test persistentvolume-controller Normal WaitForFirstConsumer waiting for first consumer to be created before binding [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test persistentvolume-controller Normal ExternalProvisioning Waiting for a volume to be created either by the external provisioner 'ebs.csi.aws.com' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal Provisioning External provisioner is provisioning volume for claim "kserve-ci-e2e-test/e2e-pvc-model-storage" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal ProvisioningSucceeded Successfully provisioned volume pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.094s (1.094s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:28 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec0c69dceeb48768325d1a53a749e65786-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.30/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedMount MountVolume.SetUp failed for volume "tls-certs" : secret "gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.60/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:46 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.60:8000/health": dial tcp 10.132.0.60:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.60:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:32 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-3c960099-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-3c960099-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:02 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-3c960099] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" in 808ms (808ms including waiting). Image size: 44914394 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.48:8001/health": dial tcp 10.132.0.48:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.49:8000/health": dial tcp 10.132.0.49:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5b86594bc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:44 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:25:52 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-50bc673d] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:07:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.44:8000/health": dial tcp 10.132.0.44:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:06:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:07:54 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-87882a8e] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 44.193s (44.193s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.62/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:40:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.62:8000/health": dial tcp 10.132.0.62:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 7f6c8cfc6 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:40:17 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test386bd5808c5c450f11fd8632e1f821fb-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:58 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-dc21cb14] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:20 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.45:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:27 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-e95b1dc1] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.42:8000/health": dial tcp 10.132.0.42:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler-67f9b884d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:01 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:19 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-7ca60146] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:06:44 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler-5444b4dbf from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:57 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:06:55 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-ba4d693a] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.55/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.54/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:25 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 686d468674 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:35 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-2577e794-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:27:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:53 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-2577e794] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:56 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.47:8000/health": dial tcp 10.132.0.47:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:51 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-59b9d263-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-59b9d263-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:25 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-59b9d263] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.50:8001/health": dial tcp 10.132.0.50:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.51:8000/health": dial tcp 10.132.0.51:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-79497db4cc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:56 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-e8706282-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:09 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-e8706282] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test80f99418f0f3eb4c80f02e6144a6d4f2-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:36 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4c1ba212] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:30 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.34:8000/health": dial tcp 10.133.0.34:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:36 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:14 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4f8c0978] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.32/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:50 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:48 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-a50492e9] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:40 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-4b931143-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-4b931143-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:21 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-4b931143] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:25 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.42:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-5b1e8f15-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-5b1e8f15-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:22 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-5b1e8f15] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-e45d1f79-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-e45d1f79-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:20 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-e45d1f79] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler-6fcb489785 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.566s (1.566s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-7dbcb75dbc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0, llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler-79866dc455 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:03 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedToRetrieveImagePullSecret Unable to retrieve some image pull secrets (llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa-dockercfg-pg2h6); attempting to pull the image may not succeed. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc44d181485fad85e662eb092f3749502f-kserve-router-scheduler-69bf65579d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler-5dd88bfbb7 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler-5f555d4d85 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler-54ccf64f4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler-6d86bd4d9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-67b4bb9646 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9, llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler-68cc9685d6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler-749449dbc8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-ccbb969bd-65ln7 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.63/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:56:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.63:8000/health": dial tcp 10.132.0.63:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 14:11:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-multiple-adapters-test-kserve-ccbb969bd-65ln7 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-multiple-adapters-test-kserve-ccbb969bd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:08 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-multiple-adapters-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-multiple-adapters-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-multiple-adapters-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:55:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:56:42 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-multiple-adapters-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.61/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:39:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-single-adapter-test-kserve-85cd6c88dc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-single-adapter-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-single-adapter-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-single-adapter-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:38:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:39:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-single-adapter-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.29/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.483s (1.483s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.29:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 3.651s (3.651s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.88s (1.88s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 2.962s (2.962s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.723s (1.723s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" in 30.447s (30.447s including waiting). Image size: 2989890188 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Liveness probe failed: timeout: failed to connect service "10.132.0.34:9003" within 1s: context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-router-scheduler-77d88bdcf4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-7c94f9d7c6 from 0 to 2 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy precise-prefix-cache-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "precise-prefix-cache-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/precise-prefix-cache-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/precise-prefix-cache-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/precise-prefix-cache-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/precise-prefix-cache-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:15 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [precise-prefix-cache-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:54:16 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-1-openshift-default-799f46c59b-fnbwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.28/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:31 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "http://10.133.0.28:15021/healthz/ready": context deadline exceeded (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:02:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning BackOff Back-off restarting failed container istio-proxy in pod router-gateway-1-openshift-default-799f46c59b-fnbwc_kserve-ci-e2e-test(f244ca92-1cf1-429e-886d-e0d585d0175f) [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:03:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.28:15021/healthz/ready": dial tcp 10.133.0.28:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-1-openshift-default-799f46c59b-fnbwc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:52:58 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-1-openshift-default-799f46c59b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:37 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-2-openshift-default-54c789bdc6-w57g9 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:19 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.44:15021/healthz/ready": dial tcp 10.133.0.44:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-2-openshift-default-54c789bdc6-w57g9 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-2-openshift-default-54c789bdc6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:17 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-2 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.56:8001/health": dial tcp 10.132.0.56:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.57/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.57:8000/health": dial tcp 10.132.0.57:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-prefill-778968b98b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-router-scheduler-5c684fd54d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-694c6ddf4c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:33 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:47 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-587fcd8566-2pqvl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:20:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:29 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-587fcd8566-2pqvl [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-router-scheduler-74545d8489 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-587fcd8566 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:18:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:19:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:05 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-7c6c58cc96-l62qh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-7c6c58cc96-l62qh [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-router-scheduler-d9ddf6c4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-7c6c58cc96 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy scheduler-inline-config-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "scheduler-inline-config-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/scheduler-inline-config-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/scheduler-inline-config-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/scheduler-inline-config-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/scheduler-inline-config-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:54:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [scheduler-inline-config-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:57:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-6hzlt to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.59/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:51 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-f2dmd to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.58/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-f2dmd [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-6hzlt [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:35:07 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy stop-feature-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "stop-feature-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-stop-feature-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/stop-feature-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/stop-feature-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/stop-feature-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:37:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [stop-feature-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.InferencePool kserve-ci-e2e-test/stop-feature-test-inference-pool [e2e-llm-inference-service] INFO e2e.llmisvc.diagnostic:diagnostic.py:56 2026-07-02 13:34:56 kserve-ci-e2e-test LLMInferenceServiceController Warning LLMInferenceServiceNotReady LLMInferenceService [stop-feature-test] is no longer Ready because of: GatewaysReady, HTTPRoutesReady, InferencePoolReady, MainWorkloadReady, PrefillWorkerWorkloadReady, PrefillWorkloadReady, RouterReady, SchedulerWorkloadReady, WorkerWorkloadReady, WorkloadsReady [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_with_lora_adapters] [2026-07-02T14:11:21.861612] end - ❌ 983.698s: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/lora-multiple-adapters-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] _ test_llm_inference_service[router-managed-workload-llmd-simulator-model-qwen2.5-0.5b] _ [e2e-llm-inference-service] [gw0] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), chunked = False [e2e-llm-inference-service] response_conn = [e2e-llm-inference-service] preload_content = False, decode_content = False, enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] # conn.request() calls http.client.*.request, not the method in [e2e-llm-inference-service] # urllib3.request. It also calls makefile (recv) on the socket. [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn.request( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] enforce_content_length=enforce_content_length, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # We are swallowing BrokenPipeError (errno.EPIPE) since the server is [e2e-llm-inference-service] # legitimately able to close the connection after sending a valid response. [e2e-llm-inference-service] # With this behaviour, the received response is still readable. [e2e-llm-inference-service] except BrokenPipeError: [e2e-llm-inference-service] pass [e2e-llm-inference-service] except OSError as e: [e2e-llm-inference-service] # MacOS/Linux [e2e-llm-inference-service] # EPROTOTYPE and ECONNRESET are needed on macOS [e2e-llm-inference-service] # https://erickt.github.io/blog/2014/11/19/adventures-in-debugging-a-potential-osx-kernel-bug/ [e2e-llm-inference-service] # Condition changed later to emit ECONNRESET instead of only EPROTOTYPE. [e2e-llm-inference-service] if e.errno != errno.EPROTOTYPE and e.errno != errno.ECONNRESET: [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset the timeout for the recv() on the socket [e2e-llm-inference-service] read_timeout = timeout_obj.read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn.is_closed: [e2e-llm-inference-service] # In Python 3 socket.py will catch EAGAIN and return None when you [e2e-llm-inference-service] # try and read into the file pointer created by http.client, which [e2e-llm-inference-service] # instead raises a BadStatusLine exception. Instead of catching [e2e-llm-inference-service] # the exception and assuming all BadStatusLine exceptions are read [e2e-llm-inference-service] # timeouts, check for a zero timeout before making the request. [e2e-llm-inference-service] if read_timeout == 0: [e2e-llm-inference-service] raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={read_timeout})" [e2e-llm-inference-service] ) [e2e-llm-inference-service] conn.timeout = read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] # Receive the response from the server [e2e-llm-inference-service] try: [e2e-llm-inference-service] > response = conn.getresponse() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:534: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def getresponse( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] ) -> HTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get the response from the server. [e2e-llm-inference-service] [e2e-llm-inference-service] If the HTTPConnection is in the correct state, returns an instance of HTTPResponse or of whatever object is returned by the response_class variable. [e2e-llm-inference-service] [e2e-llm-inference-service] If a request has not been sent or if a previous response has not be handled, ResponseNotReady is raised. If the HTTP response indicates that the connection should be closed, then it will be closed before the response is returned. When the connection is closed, the underlying socket is closed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Raise the same error as http.client.HTTPConnection [e2e-llm-inference-service] if self._response_options is None: [e2e-llm-inference-service] raise ResponseNotReady() [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset this attribute for being used again. [e2e-llm-inference-service] resp_options = self._response_options [e2e-llm-inference-service] self._response_options = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Since the connection's timeout value may have been updated [e2e-llm-inference-service] # we need to set the timeout on the socket. [e2e-llm-inference-service] self.sock.settimeout(self.timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] # This is needed here to avoid circular import errors [e2e-llm-inference-service] from .response import HTTPResponse [e2e-llm-inference-service] [e2e-llm-inference-service] # Save a reference to the shutdown function before ownership is passed [e2e-llm-inference-service] # to httplib_response [e2e-llm-inference-service] # TODO should we implement it everywhere? [e2e-llm-inference-service] _shutdown = getattr(self.sock, "shutdown", None) [e2e-llm-inference-service] [e2e-llm-inference-service] # Get the response from http.client.HTTPConnection [e2e-llm-inference-service] > httplib_response = super().getresponse() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:571: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def getresponse(self): [e2e-llm-inference-service] """Get the response from the server. [e2e-llm-inference-service] [e2e-llm-inference-service] If the HTTPConnection is in the correct state, returns an [e2e-llm-inference-service] instance of HTTPResponse or of whatever object is returned by [e2e-llm-inference-service] the response_class variable. [e2e-llm-inference-service] [e2e-llm-inference-service] If a request has not been sent or if a previous response has [e2e-llm-inference-service] not be handled, ResponseNotReady is raised. If the HTTP [e2e-llm-inference-service] response indicates that the connection should be closed, then [e2e-llm-inference-service] it will be closed before the response is returned. When the [e2e-llm-inference-service] connection is closed, the underlying socket is closed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] # if a prior response has been completed, then forget about it. [e2e-llm-inference-service] if self.__response and self.__response.isclosed(): [e2e-llm-inference-service] self.__response = None [e2e-llm-inference-service] [e2e-llm-inference-service] # if a prior response exists, then it must be completed (otherwise, we [e2e-llm-inference-service] # cannot read this response's header to determine the connection-close [e2e-llm-inference-service] # behavior) [e2e-llm-inference-service] # [e2e-llm-inference-service] # note: if a prior response existed, but was connection-close, then the [e2e-llm-inference-service] # socket and response were made independent of this HTTPConnection [e2e-llm-inference-service] # object since a new request requires that we open a whole new [e2e-llm-inference-service] # connection [e2e-llm-inference-service] # [e2e-llm-inference-service] # this means the prior response had one of two states: [e2e-llm-inference-service] # 1) will_close: this connection was reset and the prior socket and [e2e-llm-inference-service] # response operate independently [e2e-llm-inference-service] # 2) persistent: the response was retained and we await its [e2e-llm-inference-service] # isclosed() status to become true. [e2e-llm-inference-service] # [e2e-llm-inference-service] if self.__state != _CS_REQ_SENT or self.__response: [e2e-llm-inference-service] raise ResponseNotReady(self.__state) [e2e-llm-inference-service] [e2e-llm-inference-service] if self.debuglevel > 0: [e2e-llm-inference-service] response = self.response_class(self.sock, self.debuglevel, [e2e-llm-inference-service] method=self._method) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = self.response_class(self.sock, method=self._method) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > response.begin() [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:1395: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def begin(self): [e2e-llm-inference-service] if self.headers is not None: [e2e-llm-inference-service] # we've already started reading the response [e2e-llm-inference-service] return [e2e-llm-inference-service] [e2e-llm-inference-service] # read until we get a non-100 response [e2e-llm-inference-service] while True: [e2e-llm-inference-service] > version, status, reason = self._read_status() [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:325: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def _read_status(self): [e2e-llm-inference-service] > line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1") [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:286: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] b = [e2e-llm-inference-service] [e2e-llm-inference-service] def readinto(self, b): [e2e-llm-inference-service] """Read up to len(b) bytes into the writable buffer *b* and return [e2e-llm-inference-service] the number of bytes read. If the socket is non-blocking and no bytes [e2e-llm-inference-service] are available, None is returned. [e2e-llm-inference-service] [e2e-llm-inference-service] If *b* is non-empty, a 0 return value indicates that the connection [e2e-llm-inference-service] was shutdown at the other end. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self._checkClosed() [e2e-llm-inference-service] self._checkReadable() [e2e-llm-inference-service] if self._timeout_occurred: [e2e-llm-inference-service] raise OSError("cannot read from timed out object") [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return self._sock.recv_into(b) [e2e-llm-inference-service] E TimeoutError: timed out [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/socket.py:718: TimeoutError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] > response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:787: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), chunked = False [e2e-llm-inference-service] response_conn = [e2e-llm-inference-service] preload_content = False, decode_content = False, enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] # conn.request() calls http.client.*.request, not the method in [e2e-llm-inference-service] # urllib3.request. It also calls makefile (recv) on the socket. [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn.request( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] enforce_content_length=enforce_content_length, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # We are swallowing BrokenPipeError (errno.EPIPE) since the server is [e2e-llm-inference-service] # legitimately able to close the connection after sending a valid response. [e2e-llm-inference-service] # With this behaviour, the received response is still readable. [e2e-llm-inference-service] except BrokenPipeError: [e2e-llm-inference-service] pass [e2e-llm-inference-service] except OSError as e: [e2e-llm-inference-service] # MacOS/Linux [e2e-llm-inference-service] # EPROTOTYPE and ECONNRESET are needed on macOS [e2e-llm-inference-service] # https://erickt.github.io/blog/2014/11/19/adventures-in-debugging-a-potential-osx-kernel-bug/ [e2e-llm-inference-service] # Condition changed later to emit ECONNRESET instead of only EPROTOTYPE. [e2e-llm-inference-service] if e.errno != errno.EPROTOTYPE and e.errno != errno.ECONNRESET: [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset the timeout for the recv() on the socket [e2e-llm-inference-service] read_timeout = timeout_obj.read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn.is_closed: [e2e-llm-inference-service] # In Python 3 socket.py will catch EAGAIN and return None when you [e2e-llm-inference-service] # try and read into the file pointer created by http.client, which [e2e-llm-inference-service] # instead raises a BadStatusLine exception. Instead of catching [e2e-llm-inference-service] # the exception and assuming all BadStatusLine exceptions are read [e2e-llm-inference-service] # timeouts, check for a zero timeout before making the request. [e2e-llm-inference-service] if read_timeout == 0: [e2e-llm-inference-service] raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={read_timeout})" [e2e-llm-inference-service] ) [e2e-llm-inference-service] conn.timeout = read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] # Receive the response from the server [e2e-llm-inference-service] try: [e2e-llm-inference-service] response = conn.getresponse() [e2e-llm-inference-service] except (BaseSSLError, OSError) as e: [e2e-llm-inference-service] > self._raise_timeout(err=e, url=url, timeout_value=read_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:536: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] err = TimeoutError('timed out') [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] timeout_value = 60 [e2e-llm-inference-service] [e2e-llm-inference-service] def _raise_timeout( [e2e-llm-inference-service] self, [e2e-llm-inference-service] err: BaseSSLError | OSError | SocketTimeout, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] timeout_value: _TYPE_TIMEOUT | None, [e2e-llm-inference-service] ) -> None: [e2e-llm-inference-service] """Is the error actually a timeout? Will raise a ReadTimeout or pass""" [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(err, SocketTimeout): [e2e-llm-inference-service] > raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={timeout_value})" [e2e-llm-inference-service] ) from err [e2e-llm-inference-service] E urllib3.exceptions.ReadTimeoutError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:367: ReadTimeoutError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = , stream = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), verify = '/tmp/ca.crt' [e2e-llm-inference-service] cert = None, proxies = OrderedDict() [e2e-llm-inference-service] [e2e-llm-inference-service] def send( [e2e-llm-inference-service] self, request, stream=False, timeout=None, verify=True, cert=None, proxies=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Sends PreparedRequest object. Returns Response object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param request: The :class:`PreparedRequest ` being sent. [e2e-llm-inference-service] :param stream: (optional) Whether to stream the request content. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple or urllib3 Timeout object [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether [e2e-llm-inference-service] we verify the server's TLS certificate, or a string, in which case it [e2e-llm-inference-service] must be a path to a CA bundle to use [e2e-llm-inference-service] :param cert: (optional) Any user-provided SSL certificate to be trusted. [e2e-llm-inference-service] :param proxies: (optional) The proxies dictionary to apply to the request. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn = self.get_connection_with_tls_context( [e2e-llm-inference-service] request, verify, proxies=proxies, cert=cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] except LocationValueError as e: [e2e-llm-inference-service] raise InvalidURL(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] self.cert_verify(conn, request.url, verify, cert) [e2e-llm-inference-service] url = self.request_url(request, proxies) [e2e-llm-inference-service] self.add_headers( [e2e-llm-inference-service] request, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] verify=verify, [e2e-llm-inference-service] cert=cert, [e2e-llm-inference-service] proxies=proxies, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] chunked = not (request.body is None or "Content-Length" in request.headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(timeout, tuple): [e2e-llm-inference-service] try: [e2e-llm-inference-service] connect, read = timeout [e2e-llm-inference-service] timeout = TimeoutSauce(connect=connect, read=read) [e2e-llm-inference-service] except ValueError: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, " [e2e-llm-inference-service] f"or a single float to set both timeouts to the same value." [e2e-llm-inference-service] ) [e2e-llm-inference-service] elif isinstance(timeout, TimeoutSauce): [e2e-llm-inference-service] pass [e2e-llm-inference-service] else: [e2e-llm-inference-service] timeout = TimeoutSauce(connect=timeout, read=timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > resp = conn.urlopen( [e2e-llm-inference-service] method=request.method, [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] body=request.body, [e2e-llm-inference-service] headers=request.headers, [e2e-llm-inference-service] redirect=False, [e2e-llm-inference-service] assert_same_host=False, [e2e-llm-inference-service] preload_content=False, [e2e-llm-inference-service] decode_content=False, [e2e-llm-inference-service] retries=self.max_retries, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/adapters.py:667: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=7, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=6, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=5, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=4, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=3, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=2, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=1, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = RemoteDisconnected('Remote end closed connection without response') [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] body = b'{"model": "Qwen/Qwen2.5-0.5B-Instruct", "messages": [{"role": "user", "content": "What is KServe?"}], "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '119'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] > retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:841: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] response = None [e2e-llm-inference-service] error = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] _pool = [e2e-llm-inference-service] _stacktrace = [e2e-llm-inference-service] [e2e-llm-inference-service] def increment( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str | None = None, [e2e-llm-inference-service] url: str | None = None, [e2e-llm-inference-service] response: BaseHTTPResponse | None = None, [e2e-llm-inference-service] error: Exception | None = None, [e2e-llm-inference-service] _pool: ConnectionPool | None = None, [e2e-llm-inference-service] _stacktrace: TracebackType | None = None, [e2e-llm-inference-service] ) -> Self: [e2e-llm-inference-service] """Return a new Retry object with incremented retry counters. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response: A response object, or None, if the server did not [e2e-llm-inference-service] return a response. [e2e-llm-inference-service] :type response: :class:`~urllib3.response.BaseHTTPResponse` [e2e-llm-inference-service] :param Exception error: An error encountered during the request, or [e2e-llm-inference-service] None if the response was received successfully. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: A new ``Retry`` object. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if self.total is False and error: [e2e-llm-inference-service] # Disabled, indicate to re-raise the error. [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] [e2e-llm-inference-service] total = self.total [e2e-llm-inference-service] if total is not None: [e2e-llm-inference-service] total -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] connect = self.connect [e2e-llm-inference-service] read = self.read [e2e-llm-inference-service] redirect = self.redirect [e2e-llm-inference-service] status_count = self.status [e2e-llm-inference-service] other = self.other [e2e-llm-inference-service] cause = "unknown" [e2e-llm-inference-service] status = None [e2e-llm-inference-service] redirect_location = None [e2e-llm-inference-service] [e2e-llm-inference-service] if error and self._is_connection_error(error): [e2e-llm-inference-service] # Connect retry? [e2e-llm-inference-service] if connect is False: [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif connect is not None: [e2e-llm-inference-service] connect -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error and self._is_read_error(error): [e2e-llm-inference-service] # Read retry? [e2e-llm-inference-service] if read is False or method is None or not self._is_method_retryable(method): [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif read is not None: [e2e-llm-inference-service] read -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error: [e2e-llm-inference-service] # Other retry? [e2e-llm-inference-service] if other is not None: [e2e-llm-inference-service] other -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif response and response.get_redirect_location(): [e2e-llm-inference-service] # Redirect retry? [e2e-llm-inference-service] if redirect is not None: [e2e-llm-inference-service] redirect -= 1 [e2e-llm-inference-service] cause = "too many redirects" [e2e-llm-inference-service] response_redirect_location = response.get_redirect_location() [e2e-llm-inference-service] if response_redirect_location: [e2e-llm-inference-service] redirect_location = response_redirect_location [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Incrementing because of a server error like a 500 in [e2e-llm-inference-service] # status_forcelist and the given method is in the allowed_methods [e2e-llm-inference-service] cause = ResponseError.GENERIC_ERROR [e2e-llm-inference-service] if response and response.status: [e2e-llm-inference-service] if status_count is not None: [e2e-llm-inference-service] status_count -= 1 [e2e-llm-inference-service] cause = ResponseError.SPECIFIC_ERROR.format(status_code=response.status) [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] history = self.history + ( [e2e-llm-inference-service] RequestHistory(method, url, error, status, redirect_location), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] new_retry = self.new( [e2e-llm-inference-service] total=total, [e2e-llm-inference-service] connect=connect, [e2e-llm-inference-service] read=read, [e2e-llm-inference-service] redirect=redirect, [e2e-llm-inference-service] status=status_count, [e2e-llm-inference-service] other=other, [e2e-llm-inference-service] history=history, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if new_retry.is_exhausted(): [e2e-llm-inference-service] reason = error or ResponseError(cause) [e2e-llm-inference-service] > raise MaxRetryError(_pool, url, reason) from reason # type: ignore[arg-type] [e2e-llm-inference-service] E urllib3.exceptions.MaxRetryError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/util/retry.py:519: MaxRetryError [e2e-llm-inference-service] [e2e-llm-inference-service] During handling of the above exception, another exception occurred: [e2e-llm-inference-service] [e2e-llm-inference-service] def get_successful_response(): [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_case.url_getter: [e2e-llm-inference-service] service_url = test_case.url_getter(kserve_client, test_case.llm_service) [e2e-llm-inference-service] else: [e2e-llm-inference-service] service_url = get_llm_service_url(kserve_client, test_case.llm_service) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to get service URL: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] model_url = service_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] headers = {"Content-Type": "application/json"} [e2e-llm-inference-service] if extra_headers: [e2e-llm-inference-service] headers.update(extra_headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if test_case.payload_formatter is not None: [e2e-llm-inference-service] test_payload = test_case.payload_formatter(test_case) [e2e-llm-inference-service] elif test_case.prompt is not None: [e2e-llm-inference-service] test_payload = { [e2e-llm-inference-service] "model": test_case.model_name [e2e-llm-inference-service] if not extra_headers or MODEL_ROUTING_HEADER not in extra_headers [e2e-llm-inference-service] else extra_headers[MODEL_ROUTING_HEADER], [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] else: [e2e-llm-inference-service] test_payload = None [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Calling LLM service at {model_url} with payload {test_payload}") [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_payload is not None: [e2e-llm-inference-service] > response = post_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] json_data=test_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1095: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] [e2e-llm-inference-service] def post_with_retry( [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] *, [e2e-llm-inference-service] headers: Dict = None, [e2e-llm-inference-service] json_data: Union[Dict, List] = None, [e2e-llm-inference-service] data: Union[str, bytes] = None, [e2e-llm-inference-service] stream: bool = False, [e2e-llm-inference-service] timeout: float = None, [e2e-llm-inference-service] total_retries: int = DEFAULT_RETRY_TOTAL, [e2e-llm-inference-service] backoff_factor: float = DEFAULT_RETRY_BACKOFF_FACTOR, [e2e-llm-inference-service] retry_status_codes=DEFAULT_RETRY_STATUS_CODES, [e2e-llm-inference-service] ) -> requests.Response: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Send POST request with retries for transient HTTP and network failures. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if json_data is not None and data is not None: [e2e-llm-inference-service] raise ValueError("Only one of json_data or data can be provided.") [e2e-llm-inference-service] [e2e-llm-inference-service] with _retry_session( [e2e-llm-inference-service] ["POST"], total_retries, backoff_factor, retry_status_codes [e2e-llm-inference-service] ) as session: [e2e-llm-inference-service] > return session.post( [e2e-llm-inference-service] url, [e2e-llm-inference-service] json=json_data, [e2e-llm-inference-service] data=data, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] common/http_retry.py:70: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] data = None [e2e-llm-inference-service] json = {'max_tokens': 20, 'messages': [{'content': 'What is KServe?', 'role': 'user'}], 'model': 'Qwen/Qwen2.5-0.5B-Instruct'} [e2e-llm-inference-service] kwargs = {'headers': {'Content-Type': 'application/json'}, 'stream': False, 'timeout': 60} [e2e-llm-inference-service] [e2e-llm-inference-service] def post(self, url, data=None, json=None, **kwargs): [e2e-llm-inference-service] r"""Sends a POST request. Returns :class:`Response` object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: URL for the new :class:`Request` object. [e2e-llm-inference-service] :param data: (optional) Dictionary, list of tuples, bytes, or file-like [e2e-llm-inference-service] object to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param json: (optional) json to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param \*\*kwargs: Optional arguments that ``request`` takes. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.request("POST", url, data=data, json=json, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:637: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = , method = 'POST' [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions' [e2e-llm-inference-service] params = None, data = None, headers = {'Content-Type': 'application/json'} [e2e-llm-inference-service] cookies = None, files = None, auth = None, timeout = 60, allow_redirects = True [e2e-llm-inference-service] proxies = {}, hooks = None, stream = False, verify = None, cert = None [e2e-llm-inference-service] json = {'max_tokens': 20, 'messages': [{'content': 'What is KServe?', 'role': 'user'}], 'model': 'Qwen/Qwen2.5-0.5B-Instruct'} [e2e-llm-inference-service] [e2e-llm-inference-service] def request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] params=None, [e2e-llm-inference-service] data=None, [e2e-llm-inference-service] headers=None, [e2e-llm-inference-service] cookies=None, [e2e-llm-inference-service] files=None, [e2e-llm-inference-service] auth=None, [e2e-llm-inference-service] timeout=None, [e2e-llm-inference-service] allow_redirects=True, [e2e-llm-inference-service] proxies=None, [e2e-llm-inference-service] hooks=None, [e2e-llm-inference-service] stream=None, [e2e-llm-inference-service] verify=None, [e2e-llm-inference-service] cert=None, [e2e-llm-inference-service] json=None, [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Constructs a :class:`Request `, prepares it and sends it. [e2e-llm-inference-service] Returns :class:`Response ` object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: method for the new :class:`Request` object. [e2e-llm-inference-service] :param url: URL for the new :class:`Request` object. [e2e-llm-inference-service] :param params: (optional) Dictionary or bytes to be sent in the query [e2e-llm-inference-service] string for the :class:`Request`. [e2e-llm-inference-service] :param data: (optional) Dictionary, list of tuples, bytes, or file-like [e2e-llm-inference-service] object to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param json: (optional) json to send in the body of the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param headers: (optional) Dictionary of HTTP Headers to send with the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param cookies: (optional) Dict or CookieJar object to send with the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param files: (optional) Dictionary of ``'filename': file-like-objects`` [e2e-llm-inference-service] for multipart encoding upload. [e2e-llm-inference-service] :param auth: (optional) Auth tuple or callable to enable [e2e-llm-inference-service] Basic/Digest/Custom HTTP Auth. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple [e2e-llm-inference-service] :param allow_redirects: (optional) Set to True by default. [e2e-llm-inference-service] :type allow_redirects: bool [e2e-llm-inference-service] :param proxies: (optional) Dictionary mapping protocol or protocol and [e2e-llm-inference-service] hostname to the URL of the proxy. [e2e-llm-inference-service] :param hooks: (optional) Dictionary mapping hook name to one event or [e2e-llm-inference-service] list of events, event must be callable. [e2e-llm-inference-service] :param stream: (optional) whether to immediately download the response [e2e-llm-inference-service] content. Defaults to ``False``. [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether we verify [e2e-llm-inference-service] the server's TLS certificate, or a string, in which case it must be a path [e2e-llm-inference-service] to a CA bundle to use. Defaults to ``True``. When set to [e2e-llm-inference-service] ``False``, requests will accept any TLS certificate presented by [e2e-llm-inference-service] the server, and will ignore hostname mismatches and/or expired [e2e-llm-inference-service] certificates, which will make your application vulnerable to [e2e-llm-inference-service] man-in-the-middle (MitM) attacks. Setting verify to ``False`` [e2e-llm-inference-service] may be useful during local development or testing. [e2e-llm-inference-service] :param cert: (optional) if String, path to ssl client cert file (.pem). [e2e-llm-inference-service] If Tuple, ('cert', 'key') pair. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Create the Request. [e2e-llm-inference-service] req = Request( [e2e-llm-inference-service] method=method.upper(), [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] files=files, [e2e-llm-inference-service] data=data or {}, [e2e-llm-inference-service] json=json, [e2e-llm-inference-service] params=params or {}, [e2e-llm-inference-service] auth=auth, [e2e-llm-inference-service] cookies=cookies, [e2e-llm-inference-service] hooks=hooks, [e2e-llm-inference-service] ) [e2e-llm-inference-service] prep = self.prepare_request(req) [e2e-llm-inference-service] [e2e-llm-inference-service] proxies = proxies or {} [e2e-llm-inference-service] [e2e-llm-inference-service] settings = self.merge_environment_settings( [e2e-llm-inference-service] prep.url, proxies, stream, verify, cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Send the request. [e2e-llm-inference-service] send_kwargs = { [e2e-llm-inference-service] "timeout": timeout, [e2e-llm-inference-service] "allow_redirects": allow_redirects, [e2e-llm-inference-service] } [e2e-llm-inference-service] send_kwargs.update(settings) [e2e-llm-inference-service] > resp = self.send(prep, **send_kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:589: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = [e2e-llm-inference-service] kwargs = {'cert': None, 'proxies': OrderedDict(), 'stream': False, 'timeout': 60, ...} [e2e-llm-inference-service] allow_redirects = True, stream = False, hooks = {'response': []} [e2e-llm-inference-service] adapter = [e2e-llm-inference-service] start = 1783000717.0196207 [e2e-llm-inference-service] [e2e-llm-inference-service] def send(self, request, **kwargs): [e2e-llm-inference-service] """Send a given PreparedRequest. [e2e-llm-inference-service] [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Set defaults that the hooks can utilize to ensure they always have [e2e-llm-inference-service] # the correct parameters to reproduce the previous request. [e2e-llm-inference-service] kwargs.setdefault("stream", self.stream) [e2e-llm-inference-service] kwargs.setdefault("verify", self.verify) [e2e-llm-inference-service] kwargs.setdefault("cert", self.cert) [e2e-llm-inference-service] if "proxies" not in kwargs: [e2e-llm-inference-service] kwargs["proxies"] = resolve_proxies(request, self.proxies, self.trust_env) [e2e-llm-inference-service] [e2e-llm-inference-service] # It's possible that users might accidentally send a Request object. [e2e-llm-inference-service] # Guard against that specific failure case. [e2e-llm-inference-service] if isinstance(request, Request): [e2e-llm-inference-service] raise ValueError("You can only send PreparedRequests.") [e2e-llm-inference-service] [e2e-llm-inference-service] # Set up variables needed for resolve_redirects and dispatching of hooks [e2e-llm-inference-service] allow_redirects = kwargs.pop("allow_redirects", True) [e2e-llm-inference-service] stream = kwargs.get("stream") [e2e-llm-inference-service] hooks = request.hooks [e2e-llm-inference-service] [e2e-llm-inference-service] # Get the appropriate adapter to use [e2e-llm-inference-service] adapter = self.get_adapter(url=request.url) [e2e-llm-inference-service] [e2e-llm-inference-service] # Start time (approximately) of the request [e2e-llm-inference-service] start = preferred_clock() [e2e-llm-inference-service] [e2e-llm-inference-service] # Send the request [e2e-llm-inference-service] > r = adapter.send(request, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:703: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = , stream = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), verify = '/tmp/ca.crt' [e2e-llm-inference-service] cert = None, proxies = OrderedDict() [e2e-llm-inference-service] [e2e-llm-inference-service] def send( [e2e-llm-inference-service] self, request, stream=False, timeout=None, verify=True, cert=None, proxies=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Sends PreparedRequest object. Returns Response object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param request: The :class:`PreparedRequest ` being sent. [e2e-llm-inference-service] :param stream: (optional) Whether to stream the request content. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple or urllib3 Timeout object [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether [e2e-llm-inference-service] we verify the server's TLS certificate, or a string, in which case it [e2e-llm-inference-service] must be a path to a CA bundle to use [e2e-llm-inference-service] :param cert: (optional) Any user-provided SSL certificate to be trusted. [e2e-llm-inference-service] :param proxies: (optional) The proxies dictionary to apply to the request. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn = self.get_connection_with_tls_context( [e2e-llm-inference-service] request, verify, proxies=proxies, cert=cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] except LocationValueError as e: [e2e-llm-inference-service] raise InvalidURL(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] self.cert_verify(conn, request.url, verify, cert) [e2e-llm-inference-service] url = self.request_url(request, proxies) [e2e-llm-inference-service] self.add_headers( [e2e-llm-inference-service] request, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] verify=verify, [e2e-llm-inference-service] cert=cert, [e2e-llm-inference-service] proxies=proxies, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] chunked = not (request.body is None or "Content-Length" in request.headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(timeout, tuple): [e2e-llm-inference-service] try: [e2e-llm-inference-service] connect, read = timeout [e2e-llm-inference-service] timeout = TimeoutSauce(connect=connect, read=read) [e2e-llm-inference-service] except ValueError: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, " [e2e-llm-inference-service] f"or a single float to set both timeouts to the same value." [e2e-llm-inference-service] ) [e2e-llm-inference-service] elif isinstance(timeout, TimeoutSauce): [e2e-llm-inference-service] pass [e2e-llm-inference-service] else: [e2e-llm-inference-service] timeout = TimeoutSauce(connect=timeout, read=timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] resp = conn.urlopen( [e2e-llm-inference-service] method=request.method, [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] body=request.body, [e2e-llm-inference-service] headers=request.headers, [e2e-llm-inference-service] redirect=False, [e2e-llm-inference-service] assert_same_host=False, [e2e-llm-inference-service] preload_content=False, [e2e-llm-inference-service] decode_content=False, [e2e-llm-inference-service] retries=self.max_retries, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] except (ProtocolError, OSError) as err: [e2e-llm-inference-service] raise ConnectionError(err, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] except MaxRetryError as e: [e2e-llm-inference-service] if isinstance(e.reason, ConnectTimeoutError): [e2e-llm-inference-service] # TODO: Remove this in 3.0.0: see #2811 [e2e-llm-inference-service] if not isinstance(e.reason, NewConnectionError): [e2e-llm-inference-service] raise ConnectTimeout(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, ResponseError): [e2e-llm-inference-service] raise RetryError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, _ProxyError): [e2e-llm-inference-service] raise ProxyError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, _SSLError): [e2e-llm-inference-service] # This branch is for urllib3 v1.22 and later. [e2e-llm-inference-service] raise SSLError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] > raise ConnectionError(e, request=request) [e2e-llm-inference-service] E requests.exceptions.ConnectionError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/adapters.py:700: ConnectionError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator', 'model-qwen2.5-0.5b'], prompt='What is KServe?', serv... {'name': 'model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6'}]}, [e2e-llm-inference-service] 'status': None}, model_name='Qwen/Qwen2.5-0.5B-Instruct') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-gateway-ref", [e2e-llm-inference-service] "router-with-managed-route", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[0], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[0]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="custom-route-timeout-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="router-with-refs-test", [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[0], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[0]], [e2e-llm-inference-service] routes=[ROUTER_ROUTES[0], ROUTER_ROUTES[1]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=["router-managed", "workload-pd-cpu", "model-fb-opt-125m"], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="custom-route-timeout-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="router-with-refs-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[1], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[1]], [e2e-llm-inference-service] routes=[ROUTER_ROUTES[2], ROUTER_ROUTES[3]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-dp-ep-gpu", [e2e-llm-inference-service] "workload-dp-ep-prefill-gpu", [e2e-llm-inference-service] "model-deepseek-v2-lite", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="Delve into the multifaceted implications of a fully disaggregated cloud architecture, specifically " [e2e-llm-inference-service] "where the compute plane (P) and the data plane (D) are independently deployed and managed for a " [e2e-llm-inference-service] "geographically distributed, high-throughput, low-latency microservices ecosystem. Beyond the " [e2e-llm-inference-service] "fundamental challenges of network latency and data consistency, elaborate on the advanced " [e2e-llm-inference-service] "considerations and trade-offs inherent in such a setup: 1. Network Architecture and Protocols: " [e2e-llm-inference-service] "How would the network fabric and underlying protocols (e.g., RDMA, custom transport layers) need to " [e2e-llm-inference-service] "evolve to support optimal performance and minimize inter-plane communication overhead, especially for " [e2e-llm-inference-service] "synchronous operations? Discuss the role of network programmability (e.g., SDN, P4) in dynamically " [e2e-llm-inference-service] "optimizing routing and traffic flow between P and D. 2. Advanced Data Consistency and Durability: " [e2e-llm-inference-service] "Explore sophisticated data consistency models (e.g., causal consistency, strong eventual consistency) " [e2e-llm-inference-service] "and their applicability in balancing performance and data integrity across a globally distributed data plane. " [e2e-llm-inference-service] "Detail strategies for ensuring data durability and fault tolerance, including multi-region replication, " [e2e-llm-inference-service] "intelligent partitioning, and recovery mechanisms in the event of partial or full plane failures. " [e2e-llm-inference-service] "3. Dynamic Resource Orchestration and Cost Optimization: Analyze how an orchestration layer would intelligently " [e2e-llm-inference-service] "manage the independent scaling of compute (P) and data (D) resources, considering fluctuating workloads, " [e2e-llm-inference-service] "cost efficiency, and performance targets (e.g., using predictive analytics for resource provisioning). " [e2e-llm-inference-service] "Discuss mechanisms for dynamically reallocating compute nodes to different data partitions based on " [e2e-llm-inference-service] "workload patterns and data locality, potentially involving live migration strategies. " [e2e-llm-inference-service] "4. Security and Compliance in a Distributed Landscape: Address the enhanced security perimeter " [e2e-llm-inference-service] "challenges, including securing communication channels between P and D (encryption in transit, mutual TLS), " [e2e-llm-inference-service] "fine-grained access control to data at rest and in motion, and identity management across disaggregated " [e2e-llm-inference-service] "components. Discuss how such an architecture impacts compliance with regulatory frameworks (e.g., GDPR, HIPAA) " [e2e-llm-inference-service] "concerning data sovereignty, privacy, and auditability. 5. Operational Complexity and Observability: " [e2e-llm-inference-service] "Examine the increased complexity in monitoring, logging, and tracing across highly decoupled compute and " [e2e-llm-inference-service] "data planes. What specialized tooling and practices (e.g., distributed tracing with OpenTelemetry, advanced AIOps) " [e2e-llm-inference-service] "would be essential? How would incident response and troubleshooting differ in this disaggregated environment " [e2e-llm-inference-service] "compared to traditional integrated systems? Consider the challenges of pinpointing root causes across " [e2e-llm-inference-service] "independent failures. 6. Real-world Applicability and Future Trends: Identify specific industries " [e2e-llm-inference-service] "or use cases (e.g., high-frequency trading, IoT edge processing, large language model inference) " [e2e-llm-inference-service] "where the benefits of P/D disaggregation would strongly outweigh its complexities. " [e2e-llm-inference-service] "Conclude by speculating on emerging technologies or paradigms (e.g., serverless compute functions " [e2e-llm-inference-service] "directly interacting with object storage, in-memory disaggregation) that could further drive or " [e2e-llm-inference-service] "transform P/D disaggregation in cloud computing.", [e2e-llm-inference-service] max_tokens=2000, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_gpu, [e2e-llm-inference-service] pytest.mark.cluster_nvidia, [e2e-llm-inference-service] pytest.mark.cluster_nvidia_roce, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-no-scheduler", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.no_scheduler, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="This test simulates DP+EP that can run on CPU, the idea is to test the LWS-based deployment, " [e2e-llm-inference-service] "but without the resources requirements for DP+EP (GPUs and ROCe/IB).", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_multi_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler config tests [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-inline-config-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Chat completions endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] model_name="Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-configmap-ref", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-configmap-ref-test", [e2e-llm-inference-service] before_test=[create_scheduler_configmap], [e2e-llm-inference-service] after_test=[delete_scheduler_configmap], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-replicas", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-ha-replicas-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-custom-template", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-custom-template-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler v0.6 → v0.7 migration tests. [e2e-llm-inference-service] # Deploy v0.6-style configs and verify the controller migrates them [e2e-llm-inference-service] # so the v0.7 scheduler boots successfully. [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-pd-config-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-pd-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-nonzero-threshold-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-threshold-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Precise prefix KV cache routing test [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-precise-prefix-cache-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator-kvcache", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="precise-prefix-cache-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Models endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="data"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/chat/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — LoRA adapter [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] model_name=f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/models (base + LoRA) [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=assert_models_contains( [e2e-llm-inference-service] "facebook/opt-125m", [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] "lora-adapter-1", [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # PVC storage tests -- validate direct PVC volume mount with real vLLM serving [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_multi_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_inference_service(test_case: TestCase): # noqa: F811 [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] if not test_case.llm_service.metadata.annotations: [e2e-llm-inference-service] test_case.llm_service.metadata.annotations = {} [e2e-llm-inference-service] [e2e-llm-inference-service] test_case.llm_service.metadata.annotations[ [e2e-llm-inference-service] "security.opendatahub.io/enable-auth" [e2e-llm-inference-service] ] = "false" [e2e-llm-inference-service] prefix = test_case.log_prefix [e2e-llm-inference-service] [e2e-llm-inference-service] test_failed = False [e2e-llm-inference-service] try: [e2e-llm-inference-service] print(f"{prefix} Creating LLMInferenceService {service_name}") [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] print(f"{prefix} Waiting for LLMInferenceService {service_name} to be ready") [e2e-llm-inference-service] wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] print(f"{prefix} Waiting for model response from {service_name}") [e2e-llm-inference-service] > wait_for_model_response( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] test_case, [e2e-llm-inference-service] test_case.wait_timeout, [e2e-llm-inference-service] extra_headers=test_case.extra_headers, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:816: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, TestCase(base_refs=['router-managed', 'workload-llm...'name': 'model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6'}]}, [e2e-llm-inference-service] 'status': None}, model_name='Qwen/Qwen2.5-0.5B-Instruct'), 900) [e2e-llm-inference-service] kwargs = {'extra_headers': None}, func_name = 'wait_for_model_response' [e2e-llm-inference-service] timestamp_start = '2026-07-02T13:58:37.010122', start_time = 1783000717.0105004 [e2e-llm-inference-service] duration = 904.5864198207855, timestamp_end = '2026-07-02T14:13:41.596923' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-llmd-simulator', 'model-qwen2.5-0.5b'], prompt='What is KServe?', serv... {'name': 'model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6'}]}, [e2e-llm-inference-service] 'status': None}, model_name='Qwen/Qwen2.5-0.5B-Instruct') [e2e-llm-inference-service] timeout_seconds = 900, extra_headers = None [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_model_response( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] test_case: TestCase, # noqa: F811 [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] extra_headers: Optional[Dict[str, str]] = None, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def get_successful_response(): [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_case.url_getter: [e2e-llm-inference-service] service_url = test_case.url_getter(kserve_client, test_case.llm_service) [e2e-llm-inference-service] else: [e2e-llm-inference-service] service_url = get_llm_service_url(kserve_client, test_case.llm_service) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to get service URL: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] model_url = service_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] headers = {"Content-Type": "application/json"} [e2e-llm-inference-service] if extra_headers: [e2e-llm-inference-service] headers.update(extra_headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if test_case.payload_formatter is not None: [e2e-llm-inference-service] test_payload = test_case.payload_formatter(test_case) [e2e-llm-inference-service] elif test_case.prompt is not None: [e2e-llm-inference-service] test_payload = { [e2e-llm-inference-service] "model": test_case.model_name [e2e-llm-inference-service] if not extra_headers or MODEL_ROUTING_HEADER not in extra_headers [e2e-llm-inference-service] else extra_headers[MODEL_ROUTING_HEADER], [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] else: [e2e-llm-inference-service] test_payload = None [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Calling LLM service at {model_url} with payload {test_payload}") [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_payload is not None: [e2e-llm-inference-service] response = post_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] json_data=test_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = get_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] logger.error(f"❌ Failed to call model: {e}") [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to call model: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Model response is {response.status_code}: {response.text[:500]}") [e2e-llm-inference-service] [e2e-llm-inference-service] if 200 <= response.status_code < 300: [e2e-llm-inference-service] return response [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Service returned {response.status_code}: {response.text}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] > response = wait_for(get_successful_response, timeout=timeout_seconds, interval=5.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1119: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .get_successful_response at 0x7fbae4edb7e0> [e2e-llm-inference-service] timeout = 900, interval = 5.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1215: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def get_successful_response(): [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_case.url_getter: [e2e-llm-inference-service] service_url = test_case.url_getter(kserve_client, test_case.llm_service) [e2e-llm-inference-service] else: [e2e-llm-inference-service] service_url = get_llm_service_url(kserve_client, test_case.llm_service) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to get service URL: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] model_url = service_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] headers = {"Content-Type": "application/json"} [e2e-llm-inference-service] if extra_headers: [e2e-llm-inference-service] headers.update(extra_headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if test_case.payload_formatter is not None: [e2e-llm-inference-service] test_payload = test_case.payload_formatter(test_case) [e2e-llm-inference-service] elif test_case.prompt is not None: [e2e-llm-inference-service] test_payload = { [e2e-llm-inference-service] "model": test_case.model_name [e2e-llm-inference-service] if not extra_headers or MODEL_ROUTING_HEADER not in extra_headers [e2e-llm-inference-service] else extra_headers[MODEL_ROUTING_HEADER], [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] else: [e2e-llm-inference-service] test_payload = None [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Calling LLM service at {model_url} with payload {test_payload}") [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_payload is not None: [e2e-llm-inference-service] response = post_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] json_data=test_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = get_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] logger.error(f"❌ Failed to call model: {e}") [e2e-llm-inference-service] > raise AssertionError(f"❌ Failed to call model: {e}") from e [e2e-llm-inference-service] E AssertionError: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1109: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig router-managed-llmisvc-model-qw-81bfef58 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig router-managed-llmisvc-model-qw-81bfef58 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig router-managed-llmisvc-model-qw-81bfef58 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig workload-llmd-simulator-llmisvc-947004c3 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig workload-llmd-simulator-llmisvc-947004c3 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig workload-llmd-simulator-llmisvc-947004c3 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_inference_service] [2026-07-02T13:58:00.248703] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-llmd-simulator', 'model-qwen2.5-0.5b'], prompt='What is KServe?', service_name='llmisvc-model-qwen2-5-0-5b-rout-4c1ba212', endpoint='/v1/chat/completions', max_tokens=20, payload_formatter=, response_assertion=.response_assertion at 0x7fbae5307920>, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': None, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-qwen2-5-0-5b-rout-4c1ba212', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-qw-81bfef58'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-llmisvc-947004c3'}, [e2e-llm-inference-service] {'name': 'model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6'}]}, [e2e-llm-inference-service] 'status': None}, model_name='Qwen/Qwen2.5-0.5B-Instruct')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-02T13:58:00.261307] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-qwen2-5-0-5b-rout-4c1ba212', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-qw-81bfef58'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-llmisvc-947004c3'}, [e2e-llm-inference-service] {'name': 'model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-02T13:58:00.346382] end - ✅ in 0.085s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-02T13:58:00.346507] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-qwen2-5-0-5b-rout-4c1ba212', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-qw-81bfef58'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-llmisvc-947004c3'}, [e2e-llm-inference-service] {'name': 'model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'Inference Pool kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:58:12Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: "False" (reason "BackendNotFound", message "backend(llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference--ip-9e4c11fc.kserve-ci-e2e-test.svc.cluster.local) not found")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:58:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:58:12Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: "False" (reason "BackendNotFound", message "backend(llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference--ip-9e4c11fc.kserve-ci-e2e-test.svc.cluster.local) not found")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:58:12Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: "False" (reason "BackendNotFound", message "backend(llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference--ip-9e4c11fc.kserve-ci-e2e-test.svc.cluster.local) not found")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:58:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:58:13Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:58:13Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:58:13Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:58:13Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:58:13Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T13:58:13Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T13:58:12Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T13:58:13Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:58:04Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T13:58:13Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T13:58:13Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T13:58:12Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T13:58:13Z', 'status': 'True', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [wait_for_llm_isvc_ready] [2026-07-02T13:58:37.009936] end - ✅ in 36.663s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_model_response] [2026-07-02T13:58:37.010122] start - args=(, TestCase(base_refs=['router-managed', 'workload-llmd-simulator', 'model-qwen2.5-0.5b'], prompt='What is KServe?', service_name='llmisvc-model-qwen2-5-0-5b-rout-4c1ba212', endpoint='/v1/chat/completions', max_tokens=20, payload_formatter=, response_assertion=.response_assertion at 0x7fbae5307920>, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-qwen2-5-0-5b-rout-4c1ba212', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-qw-81bfef58'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-llmisvc-947004c3'}, [e2e-llm-inference-service] {'name': 'model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6'}]}, [e2e-llm-inference-service] 'status': None}, model_name='Qwen/Qwen2.5-0.5B-Instruct'), 900), kwargs={'extra_headers': None} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [get_llm_service_url] [2026-07-02T13:58:37.010505] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'llmisvc-model-qwen2-5-0-5b-rout-4c1ba212', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-llmisvc-model-qw-81bfef58'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-llmisvc-947004c3'}, [e2e-llm-inference-service] {'name': 'model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [get_llm_service_url] [2026-07-02T13:58:37.018273] end - ✅ in 0.008s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1092 Calling LLM service at http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions with payload {'model': 'Qwen/Qwen2.5-0.5B-Instruct', 'messages': [{'role': 'user', 'content': 'What is KServe?'}], 'max_tokens': 20} [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=7, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=6, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=5, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=4, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=3, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=2, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=1, connect=None, read=None, redirect=None, status=None)) after connection broken by 'RemoteDisconnected('Remote end closed connection without response')': /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=0, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1108 ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1219 Timed out waiting: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_model_response] [2026-07-02T14:13:41.596923] end - ❌ 904.586s: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:831 [router-managed-workload-llmd-simulator-model-qwen2.5-0.5b] ❌ ERROR: Failed to call llm inference service llmisvc-model-qwen2-5-0-5b-rout-4c1ba212: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1240 🔍 # Diagnostics for 'llmisvc-model-qwen2-5-0-5b-rout-4c1ba212' in 'kserve-ci-e2e-test' [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1241 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1242 # LLMInferenceService llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1245 apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] security.opendatahub.io/enable-auth: 'false' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:00Z' [e2e-llm-inference-service] finalizers: [e2e-llm-inference-service] - serving.kserve.io/llmisvc-finalizer [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:security.opendatahub.io/enable-auth: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:baseRefs: {} [e2e-llm-inference-service] manager: OpenAPI-Generator [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:00Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:finalizers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] v:"serving.kserve.io/llmisvc-finalizer": {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:00Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:addresses: {} [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-router-route: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-scheduler: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-tracing: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-worker-data-parallel: {} [e2e-llm-inference-service] f:appliedConfigs: {} [e2e-llm-inference-service] f:conditions: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:router: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:gateways: {} [e2e-llm-inference-service] f:scheduler: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:inferencePool: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:service: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:url: {} [e2e-llm-inference-service] f:workloads: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:primary: {} [e2e-llm-inference-service] f:scheduler: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] resourceVersion: '72093' [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] baseRefs: [e2e-llm-inference-service] - name: router-managed-llmisvc-model-qw-81bfef58 [e2e-llm-inference-service] - name: workload-llmd-simulator-llmisvc-947004c3 [e2e-llm-inference-service] - name: model-qwen2-5-0-5b-llmisvc-mode-8ec1f6c6 [e2e-llm-inference-service] model: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uri: '' [e2e-llm-inference-service] status: [e2e-llm-inference-service] addresses: [e2e-llm-inference-service] - name: gateway-external-model-routing [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/ [e2e-llm-inference-service] - name: gateway-external [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] - name: gateway-internal-model-routing [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/ [e2e-llm-inference-service] - name: gateway-internal [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-template: kserve-config-llm-decode-template [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-worker-data-parallel: kserve-config-llm-decode-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-template: kserve-config-llm-prefill-template [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-worker-data-parallel: kserve-config-llm-prefill-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-router-route: kserve-config-llm-router-route [e2e-llm-inference-service] serving.kserve.io/config-llm-scheduler: kserve-config-llm-scheduler [e2e-llm-inference-service] serving.kserve.io/config-llm-template: kserve-config-llm-template [e2e-llm-inference-service] serving.kserve.io/config-llm-tracing: kserve-config-llm-tracing [e2e-llm-inference-service] serving.kserve.io/config-llm-worker-data-parallel: kserve-config-llm-worker-data-parallel [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: HTTPRoutesReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:12Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: InferencePoolReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: MainWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:04Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: PresetsCombined [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Ready [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: RouterReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: SchedulerWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: WorkloadsReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:44 TIME NAMESPACE SOURCE TYPE REASON MESSAGE [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:45 -------------------------------------------------------------------------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-6599bc8fd-8sr69 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-6599bc8fd-8sr69 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-router-scheduler-6fcbb55f87 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-6599bc8fd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:52 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy auth-disabled-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "auth-disabled-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-disabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-disabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-disabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-disabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:56 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-disabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-67c57d579b-gtwnj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:05 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 27.649s (27.649s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.35:8000/health": dial tcp 10.132.0.35:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-67c57d579b-gtwnj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.31/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.181s (1.181s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-router-scheduler-688c79755d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-67c57d579b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-enabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-enabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-enabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-enabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-enabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-enabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-6cd6c5884-rrbrr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-6cd6c5884-rrbrr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-router-scheduler-5b958c76d8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-6cd6c5884 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-invalid-token-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-invalid-token-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-invalid-token-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-invalid-token-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-invalid-token-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:24 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-invalid-token-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:25 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.52:8001/health": dial tcp 10.132.0.52:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.53/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.53:8000/health": dial tcp 10.132.0.53:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-router-scheduler-f7968dfdd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-79d6b9cc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:31 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-pd-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-pd-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:43 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-98989556b-vl89r to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:14:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.45:8000/health": dial tcp 10.132.0.45:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-98989556b-vl89r [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-router-scheduler-579b7545d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-98989556b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:57 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:15:01 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/e2e-pvc-model-download-4s522 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:26 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test job-controller Normal SuccessfulCreate Created pod: e2e-pvc-model-download-4s522 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:36 kserve-ci-e2e-test job-controller Normal Completed Job completed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:20 kserve-ci-e2e-test persistentvolume-controller Normal WaitForFirstConsumer waiting for first consumer to be created before binding [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test persistentvolume-controller Normal ExternalProvisioning Waiting for a volume to be created either by the external provisioner 'ebs.csi.aws.com' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal Provisioning External provisioner is provisioning volume for claim "kserve-ci-e2e-test/e2e-pvc-model-storage" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal ProvisioningSucceeded Successfully provisioned volume pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.094s (1.094s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:28 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec0c69dceeb48768325d1a53a749e65786-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.30/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedMount MountVolume.SetUp failed for volume "tls-certs" : secret "gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.60/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:46 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.60:8000/health": dial tcp 10.132.0.60:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.60:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:32 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-3c960099-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-3c960099-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:02 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-3c960099] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" in 808ms (808ms including waiting). Image size: 44914394 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.48:8001/health": dial tcp 10.132.0.48:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.49:8000/health": dial tcp 10.132.0.49:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5b86594bc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:44 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:52 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-50bc673d] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:07:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.44:8000/health": dial tcp 10.132.0.44:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:07:54 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-87882a8e] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 44.193s (44.193s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.62/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:40:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.62:8000/health": dial tcp 10.132.0.62:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 7f6c8cfc6 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:40:17 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test386bd5808c5c450f11fd8632e1f821fb-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:58 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-dc21cb14] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:20 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.45:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:27 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-e95b1dc1] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.42:8000/health": dial tcp 10.132.0.42:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler-67f9b884d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:01 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:19 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-7ca60146] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:44 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler-5444b4dbf from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:57 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:55 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-ba4d693a] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.55/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.54/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:25 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 686d468674 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:35 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-2577e794-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:53 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-2577e794] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:56 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.47:8000/health": dial tcp 10.132.0.47:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:51 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-59b9d263-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-59b9d263-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:25 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-59b9d263] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.50:8001/health": dial tcp 10.132.0.50:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.51:8000/health": dial tcp 10.132.0.51:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-79497db4cc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:56 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-e8706282-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:09 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-e8706282] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test80f99418f0f3eb4c80f02e6144a6d4f2-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:36 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4c1ba212] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:30 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.34:8000/health": dial tcp 10.133.0.34:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:36 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:14 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4f8c0978] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.32/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:50 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:48 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-a50492e9] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:40 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-4b931143-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-4b931143-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:21 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-4b931143] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:25 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.42:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-5b1e8f15-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-5b1e8f15-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:22 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-5b1e8f15] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-e45d1f79-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-e45d1f79-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:20 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-e45d1f79] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler-6fcb489785 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.566s (1.566s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-7dbcb75dbc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0, llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler-79866dc455 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:03 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedToRetrieveImagePullSecret Unable to retrieve some image pull secrets (llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa-dockercfg-pg2h6); attempting to pull the image may not succeed. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc44d181485fad85e662eb092f3749502f-kserve-router-scheduler-69bf65579d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler-5dd88bfbb7 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler-5f555d4d85 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler-54ccf64f4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler-6d86bd4d9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-67b4bb9646 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9, llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler-68cc9685d6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler-749449dbc8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-ccbb969bd-65ln7 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.63/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:56:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.63:8000/health": dial tcp 10.132.0.63:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-multiple-adapters-test-kserve-ccbb969bd-65ln7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-multiple-adapters-test-kserve-ccbb969bd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:08 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-multiple-adapters-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-multiple-adapters-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-multiple-adapters-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:56:42 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-multiple-adapters-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.61/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-single-adapter-test-kserve-85cd6c88dc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-single-adapter-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-single-adapter-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-single-adapter-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-single-adapter-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.29/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.483s (1.483s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.29:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 3.651s (3.651s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.88s (1.88s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 2.962s (2.962s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.723s (1.723s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" in 30.447s (30.447s including waiting). Image size: 2989890188 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Liveness probe failed: timeout: failed to connect service "10.132.0.34:9003" within 1s: context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-router-scheduler-77d88bdcf4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-7c94f9d7c6 from 0 to 2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy precise-prefix-cache-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "precise-prefix-cache-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/precise-prefix-cache-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/precise-prefix-cache-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/precise-prefix-cache-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/precise-prefix-cache-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:15 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [precise-prefix-cache-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:16 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-1-openshift-default-799f46c59b-fnbwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.28/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:31 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "http://10.133.0.28:15021/healthz/ready": context deadline exceeded (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning BackOff Back-off restarting failed container istio-proxy in pod router-gateway-1-openshift-default-799f46c59b-fnbwc_kserve-ci-e2e-test(f244ca92-1cf1-429e-886d-e0d585d0175f) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.28:15021/healthz/ready": dial tcp 10.133.0.28:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-1-openshift-default-799f46c59b-fnbwc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:58 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-1-openshift-default-799f46c59b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:37 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-2-openshift-default-54c789bdc6-w57g9 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:19 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.44:15021/healthz/ready": dial tcp 10.133.0.44:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-2-openshift-default-54c789bdc6-w57g9 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-2-openshift-default-54c789bdc6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:17 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.56:8001/health": dial tcp 10.132.0.56:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.57/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.57:8000/health": dial tcp 10.132.0.57:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-prefill-778968b98b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-router-scheduler-5c684fd54d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-694c6ddf4c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:33 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:47 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-587fcd8566-2pqvl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:20:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:29 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-587fcd8566-2pqvl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-router-scheduler-74545d8489 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-587fcd8566 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:05 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-7c6c58cc96-l62qh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-7c6c58cc96-l62qh [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-router-scheduler-d9ddf6c4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-7c6c58cc96 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy scheduler-inline-config-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "scheduler-inline-config-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/scheduler-inline-config-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/scheduler-inline-config-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/scheduler-inline-config-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/scheduler-inline-config-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [scheduler-inline-config-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-6hzlt to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.59/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:51 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-f2dmd to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.58/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-f2dmd [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-6hzlt [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:07 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy stop-feature-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "stop-feature-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-stop-feature-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/stop-feature-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/stop-feature-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/stop-feature-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [stop-feature-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.InferencePool kserve-ci-e2e-test/stop-feature-test-inference-pool [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:56 kserve-ci-e2e-test LLMInferenceServiceController Warning LLMInferenceServiceNotReady LLMInferenceService [stop-feature-test] is no longer Ready because of: GatewaysReady, HTTPRoutesReady, InferencePoolReady, MainWorkloadReady, PrefillWorkerWorkloadReady, PrefillWorkloadReady, RouterReady, SchedulerWorkloadReady, WorkerWorkloadReady, WorkloadsReady [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/tls-verification-test-kserve-6d6946dcd6-bdddw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.64/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.64:8000/health": dial tcp 10.132.0.64:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: tls-verification-test-kserve-6d6946dcd6-bdddw [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.65/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set tls-verification-test-kserve-router-scheduler-7f955879f5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set tls-verification-test-kserve-6d6946dcd6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:24 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy tls-verification-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/tls-verification-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "tls-verification-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/tls-verification-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/tls-verification-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-tls-verification-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/tls-verification-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/tls-verification-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/tls-verification-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/tls-verification-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/tls-verification-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [tls-verification-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 I0702 13:58:02.890315 1 config.go:602] "Configuration:" =< [e2e-llm-inference-service] { [e2e-llm-inference-service] "IP": "", [e2e-llm-inference-service] "PodName": "", [e2e-llm-inference-service] "PodNameSpace": "", [e2e-llm-inference-service] "VllmDevMode": false, [e2e-llm-inference-service] "block-size": 16, [e2e-llm-inference-service] "data-parallel-rank": -1, [e2e-llm-inference-service] "data-parallel-size": 1, [e2e-llm-inference-service] "dataset-in-memory": false, [e2e-llm-inference-service] "dataset-path": "", [e2e-llm-inference-service] "dataset-table-name": "llmd", [e2e-llm-inference-service] "dataset-url": "", [e2e-llm-inference-service] "default-embedding-dimensions": 384, [e2e-llm-inference-service] "ec-transfer-config": "", [e2e-llm-inference-service] "enable-kvcache": false, [e2e-llm-inference-service] "enable-prefix-caching": false, [e2e-llm-inference-service] "enable-request-id-headers": false, [e2e-llm-inference-service] "enable-sleep-mode": false, [e2e-llm-inference-service] "enforce-eager": false, [e2e-llm-inference-service] "event-batch-size": 16, [e2e-llm-inference-service] "failure-injection-rate": 0, [e2e-llm-inference-service] "failure-types": null, [e2e-llm-inference-service] "fake-metrics": null, [e2e-llm-inference-service] "fake-metrics-refresh-interval": 100000000, [e2e-llm-inference-service] "global-cache-hit-threshold": 0, [e2e-llm-inference-service] "hash-seed": "", [e2e-llm-inference-service] "inter-token-latency": 0, [e2e-llm-inference-service] "inter-token-latency-std-dev": 0, [e2e-llm-inference-service] "kv-cache-size": 1024, [e2e-llm-inference-service] "kv-cache-transfer-latency": 0, [e2e-llm-inference-service] "kv-cache-transfer-latency-std-dev": 0, [e2e-llm-inference-service] "kv-cache-transfer-time-per-token": 0, [e2e-llm-inference-service] "kv-cache-transfer-time-std-dev": 0, [e2e-llm-inference-service] "latency-calculator": "", [e2e-llm-inference-service] "lora-modules": null, [e2e-llm-inference-service] "max-cpu-loras": 1, [e2e-llm-inference-service] "max-loras": 1, [e2e-llm-inference-service] "max-model-len": 1024, [e2e-llm-inference-service] "max-num-seqs": 5, [e2e-llm-inference-service] "max-tool-call-array-param-length": 5, [e2e-llm-inference-service] "max-tool-call-integer-param": 100, [e2e-llm-inference-service] "max-tool-call-number-param": 100, [e2e-llm-inference-service] "max-waiting-queue-length": 1000, [e2e-llm-inference-service] "min-tool-call-array-param-length": 1, [e2e-llm-inference-service] "min-tool-call-integer-param": 0, [e2e-llm-inference-service] "min-tool-call-number-param": 0, [e2e-llm-inference-service] "mm-encoder-only": false, [e2e-llm-inference-service] "mm-processor-kwargs": "", [e2e-llm-inference-service] "mode": "random", [e2e-llm-inference-service] "model": "Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] "object-tool-call-not-required-field-probability": 50, [e2e-llm-inference-service] "port": 8000, [e2e-llm-inference-service] "prefill-overhead": 0, [e2e-llm-inference-service] "prefill-time-per-token": 0, [e2e-llm-inference-service] "prefill-time-std-dev": 0, [e2e-llm-inference-service] "seed": 1783000682889895400, [e2e-llm-inference-service] "self-signed-certs": false, [e2e-llm-inference-service] "served-model-name": [ [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ], [e2e-llm-inference-service] "ssl-certfile": "/var/run/kserve/tls/tls.crt", [e2e-llm-inference-service] "ssl-keyfile": "/var/run/kserve/tls/tls.key", [e2e-llm-inference-service] "time-factor-under-load": 1, [e2e-llm-inference-service] "time-to-first-token": 0, [e2e-llm-inference-service] "time-to-first-token-std-dev": 0, [e2e-llm-inference-service] "tool-call-not-required-param-probability": 50, [e2e-llm-inference-service] "uds-socket-path": "/tmp/tokenizer/tokenizer-uds.socket", [e2e-llm-inference-service] "zmq-endpoint": "tcp://127.0.0.1:5557" [e2e-llm-inference-service] } [e2e-llm-inference-service] > [e2e-llm-inference-service] I0702 13:58:02.928825 1 tokenizer.go:104] "Model is not a real HF model, using simulated tokenizer" model="Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] I0702 13:58:02.933557 1 context.go:138] "No dataset path or URL provided, using random text for responses" [e2e-llm-inference-service] I0702 13:58:02.933628 1 communication.go:49] "Starting communication layer" [e2e-llm-inference-service] I0702 13:58:02.933667 1 simulator.go:188] "Start processing routine" [e2e-llm-inference-service] I0702 13:58:02.933948 1 http_server_tls.go:44] "HTTPS server starting with certificate files" cert="/var/run/kserve/tls/tls.crt" key="/var/run/kserve/tls/tls.key" [e2e-llm-inference-service] I0702 13:58:02.934018 1 grpc.go:126] "Server starting" protocol="gRPC" port=8000 [e2e-llm-inference-service] I0702 13:58:02.935011 1 http.go:96] "Server starting" protocol="HTTPS" port=8000 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 {"level":"info","ts":1783000684.1554873,"logger":"setup","caller":"runner/runner.go:196","msg":"GIE build","commit-sha":"181aa8358916e19b8844ccc752b2d6153d4b2ad6","build-ref":"v0.9.0-rc.2"} [e2e-llm-inference-service] Flag --model-server-metrics-scheme has been deprecated, This flag is deprecated. Configure via EndpointPickerConfig data layer plugin parameters instead. [e2e-llm-inference-service] {"level":"info","ts":1783000684.1557543,"logger":"setup","caller":"runner/runner.go:217","msg":"Flags processed","flags":{"cert-path":"/var/run/kserve/tls","config-file":"","config-text":"apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type: prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n- name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n","disable-endpoint-subset-filter":false,"enable-cert-reload":true,"enable-grpc-stream-metrics":false,"enable-pprof":true,"endpoint-selector":"","endpoint-target-ports":{},"grpc-health-port":9003,"grpc-max-recv-msg-size":"","grpc-max-send-msg-size":"","grpc-port":9002,"ha-enable-leader-election":false,"health-checking":false,"metrics-endpoint-auth":true,"metrics-port":9090,"metrics-staleness-threshold":2000000000,"model-server-metrics-https-insecure-skip-verify":true,"model-server-metrics-path":"/metrics","model-server-metrics-port":0,"model-server-metrics-scheme":"https","pool-group":"inference.networking.k8s.io","pool-name":"llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool","pool-namespace":"kserve-ci-e2e-test","refresh-metrics-interval":50000000,"refresh-prometheus-metrics-interval":5000000000,"secure-serving":true,"tracing":true,"v":2,"zap-devel":{},"zap-encoder":{},"zap-log-level":{},"zap-stacktrace-level":{},"zap-time-encoding":{}}} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1564596,"logger":"setup.trace","caller":"tracing/telemetry.go:123","msg":"init OTel trace exporter","type":"console"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1572437,"caller":"loader/configloader.go:89","msg":"DEPRECATION: apiVersion inference.networking.x-k8s.io/v1alpha1/EndpointPickerConfig is deprecated","replacement":"llm-d.ai/v1alpha1/EndpointPickerConfig"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1572857,"caller":"loader/configloader.go:121","msg":"Loaded raw configuration","config":"{Plugins: [{Type: single-profile-handler} {Type: queue-scorer} {Type: prefix-cache-scorer} {Type: max-score-picker}], SchedulingProfiles: [{Name: default, Plugins: [{PluginRef: queue-scorer, Weight: 2.00} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: max-score-picker}]}]}"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1572955,"logger":"setup","caller":"runner/runner.go:622","msg":"Data layer: ENABLED"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1575518,"logger":"setup","caller":"runner/runner.go:281","msg":"Raw config after phase one","config":{"apiVersion":"inference.networking.x-k8s.io/v1alpha1","dataLayer":null,"kind":"EndpointPickerConfig","plugins":[{"name":"single-profile-handler","parameters":null,"type":"single-profile-handler"},{"name":"queue-scorer","parameters":null,"type":"queue-scorer"},{"name":"prefix-cache-scorer","parameters":null,"type":"prefix-cache-scorer"},{"name":"max-score-picker","parameters":null,"type":"max-score-picker"}],"schedulingProfiles":[{"name":"default","plugins":[{"pluginRef":"queue-scorer","weight":2},{"pluginRef":"prefix-cache-scorer","weight":3},{"pluginRef":"max-score-picker","weight":null}]}]}} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1708703,"logger":"utilization-detector/utilization-detector","caller":"utilization/detector.go:83","msg":"Creating new UtilizationDetector","queueDepthThreshold":5,"kvCacheUtilThreshold":0.8,"metricsStalenessThreshold":"200ms","headroom":0} [e2e-llm-inference-service] {"level":"info","ts":1783000684.170954,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"vllm","mapping":"Mapping{all specs enabled}"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1709957,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"sglang","mapping":"Mapping{disabled: [lora]}"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1710308,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"trtllm-serve","mapping":"Mapping{disabled: [lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1710894,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"triton-tensorrt-llm","mapping":"Mapping{disabled: [lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.171109,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"triton","mapping":"Mapping{disabled: [kv, lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1711524,"caller":"loader/configloader.go:154","msg":"Instantiated all plugins and applied system defaults. Effective raw configuration","config":"{Plugins: [{Name: single-profile-handler, Type: single-profile-handler} {Name: queue-scorer, Type: queue-scorer} {Name: prefix-cache-scorer, Type: prefix-cache-scorer} {Name: max-score-picker, Type: max-score-picker} {Name: fcfs-ordering-policy, Type: fcfs-ordering-policy} {Name: global-strict-fairness-policy, Type: global-strict-fairness-policy} {Name: static-usage-limit-policy, Type: static-usage-limit-policy} {Name: openai-parser, Type: openai-parser} {Name: anthropic-parser, Type: anthropic-parser} {Name: vllmhttp-parser, Type: vllmhttp-parser} {Name: utilization-detector, Type: utilization-detector} {Name: metrics-data-source, Type: metrics-data-source} {Name: core-metrics-extractor, Type: core-metrics-extractor}], SchedulingProfiles: [{Name: default, Plugins: [{PluginRef: queue-scorer, Weight: 2.00} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: max-score-picker}]}], DataLayer: {Sources: [{PluginRef: metrics-data-source, Extractors: [{PluginRef: core-metrics-extractor}]}], Discovery: }, FlowControl: {MaxBytes: unlimited, MaxRequests: unlimited, SaturationDetector: {PluginRef: utilization-detector}}, RequestHandler: {Parsers: [{PluginRef: openai-parser}, {PluginRef: anthropic-parser}, {PluginRef: vllmhttp-parser}]}}"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.171218,"caller":"approximateprefix/plugin.go:88","msg":"Prefix DataProducer initialized","config":{"autoTune":true,"blockSizeTokens":16,"blockSize":0,"maxPrefixBlocksToMatch":2048,"maxPrefixTokensToMatch":131072,"lruCapacityPerServer":31250}} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1712863,"caller":"approximateprefix/plugin.go:111","msg":"WARNING: configured blockSizeTokens is below the recommended minimum, overriding it.","blockSizeTokens":16,"minimum":64,"issue":"https://github.com/llm-d/llm-d-router/issues/1158"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1713207,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"approx-prefix-cache-producer/approx-prefix-cache-producer","dataKey":"PrefixCacheMatchInfoDataKey/approx-prefix-cache-producer","consumer":"prefix-cache-scorer"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1713507,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"token-producer/token-producer","dataKey":"TokenizedPrompt/token-producer","consumer":"approx-prefix-cache-producer"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1714306,"caller":"runner/runner.go:685","msg":"loaded configuration from file/text successfully"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1714406,"logger":"setup","caller":"runner/runner.go:308","msg":"EPP config after phase two","config":"{SchedulerConfig:{ProfileHandler: single-profile-handler/single-profile-handler, Profiles: map[default:{Filters: [], Scorers: [queue-scorer/queue-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000], Picker: max-score-picker/max-score-picker}]} SaturationDetector:0xc000711dc0 DataConfig:{Sources:[{Plugin:0xc000153560 Extractors:[0xc00040ab40]}]} FlowControlConfig: ParserRegistry:0xc0002a5500}"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1864343,"logger":"setup","caller":"runner/runner.go:352","msg":"Setting pprof handlers"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1864607,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/block"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.186475,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/profile"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1864798,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/symbol"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.186485,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/allocs"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1864893,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/threadcreate"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1864936,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/mutex"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1864984,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1865027,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/cmdline"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1865096,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/trace"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.186514,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/heap"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1865184,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/goroutine"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1865256,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/plugins/state"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1865308,"logger":"setup","caller":"runner/runner.go:373","msg":"parsed config","scheduler-config":"{ProfileHandler: single-profile-handler/single-profile-handler, Profiles: map[default:{Filters: [], Scorers: [queue-scorer/queue-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000], Picker: max-score-picker/max-score-picker}]}"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.186555,"logger":"setup","caller":"datalayer/runtime.go:99","msg":"Configuring datalayer runtime","numSources":1} [e2e-llm-inference-service] {"level":"info","ts":1783000684.186563,"logger":"setup","caller":"datalayer/runtime.go:118","msg":"Processing source","source":"metrics-data-source","numExtractors":1} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1865783,"logger":"setup","caller":"datalayer/runtime.go:147","msg":"Source configured","source":"metrics-data-source","extractors":["core-metrics-extractor/core-metrics-extractor"]} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1865914,"logger":"setup","caller":"datalayer/runtime.go:206","msg":"Datalayer runtime configured","pollers":1,"notifiers":0,"endpointSources":0} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1866002,"logger":"setup","caller":"runner/runner.go:833","msg":"Experimental Flow Control layer is disabled, using legacy admission control"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1866858,"logger":"setup","caller":"runner/runner.go:721","msg":"ExtProc server runner added to manager."} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1866953,"logger":"setup","caller":"runner/runner.go:260","msg":"Controller manager starting"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1867232,"logger":"controller-runtime.metrics","caller":"server/server.go:208","msg":"Starting metrics server"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1869256,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"health"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1871862,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"health","port":9003} [e2e-llm-inference-service] {"level":"info","ts":1783000684.187218,"logger":"controller-runtime.metrics","caller":"server/server.go:247","msg":"Serving metrics server","bindAddress":":9090","secure":false} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1872644,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","source":"kind source: *v1.InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.187343,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite","source":"kind source: *v1alpha2.InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1873553,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective","source":"kind source: *v1alpha2.InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.187386,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"pod","controllerGroup":"","controllerKind":"Pod","source":"kind source: *v1.Pod"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.187843,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"ext-proc"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.187919,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"ext-proc","port":9002} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1913223,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceObjective","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.1913366,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceModelRewrite","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.19152,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.InferencePool","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.192337,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.Pod","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.2885184,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.2887387,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783000684.288696,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"pod","controllerGroup":"","controllerKind":"Pod"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.2887807,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"pod","controllerGroup":"","controllerKind":"Pod","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783000684.2887092,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.2888517,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783000684.3895047,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1783000684.3895643,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783000684.3897111,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool","reconcileID":"121f99d3-5ebc-4bbe-9eae-071603b24626"} [e2e-llm-inference-service] {"level":"info","ts":1783000690.710773,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool","reconcileID":"3980185f-0d7d-4b59-99df-15a0cd2fb2aa"} [e2e-llm-inference-service] {"level":"info","ts":1783000693.046885,"caller":"controller/pod_reconciler.go:99","msg":"Pod already exists","controller":"pod","controllerGroup":"","controllerKind":"Pod","Pod":{"name":"llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z","reconcileID":"54558a02-208c-442e-b949-aeb91eaaa876"} [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 9c06c4c9-fa0f-4753-93e6-19ca799d1951 [e2e-llm-inference-service] resourceVersion: '72083' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T13:58:35Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.134.0.52 [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l [e2e-llm-inference-service] uid: da26fdce-65db-48f8-9d58-b72f0700b8c8 [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 91fb3e3b-8e8b-45c3-8364-95ae711cd3e0 [e2e-llm-inference-service] resourceVersion: '71797' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.134.0.51 [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z [e2e-llm-inference-service] uid: 2b688cb9-a351-4747-aa9f-01cee48d1c3d [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z [e2e-llm-inference-service] generateName: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff4- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 2b688cb9-a351-4747-aa9f-01cee48d1c3d [e2e-llm-inference-service] resourceVersion: '71794' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 797dfc7ff4 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.134.0.51/23"],"mac_address":"0a:58:0a:86:00:33","gateway_ips":["10.134.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.134.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.134.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.134.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.134.0.1"}],"ip_address":"10.134.0.51/23","gateway_ip":"10.134.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.134.0.51\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:86:00:33\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff4 [e2e-llm-inference-service] uid: cafdc8c6-2fc7-4e63-8b18-19606b53c8e4 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-138-159 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"cafdc8c6-2fc7-4e63-8b18-19606b53c8e4"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.134.0.51"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kube-api-access-264zb [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/llm-d-inference-sim [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --model [e2e-llm-inference-service] - Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] - --mode [e2e-llm-inference-service] - random [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kube-api-access-264zb [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: default [e2e-llm-inference-service] serviceAccount: default [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] fsGroup: 1000690000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] hostIP: 10.0.138.159 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.138.159 [e2e-llm-inference-service] podIP: 10.134.0.51 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.134.0.51 [e2e-llm-inference-service] startTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] imageID: ghcr.io/llm-d/llm-d-inference-sim@sha256:bab162bd25e2ed8b15022387cdb223023aeb33be49476af9f0115c0398fb8ff5 [e2e-llm-inference-service] containerID: cri-o://9074c47218029a2d40daf10467d8df61f2b018101ae478c5b3d43c64146dac85 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-264zb [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l [e2e-llm-inference-service] generateName: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler-79866dc455- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: da26fdce-65db-48f8-9d58-b72f0700b8c8 [e2e-llm-inference-service] resourceVersion: '72082' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 79866dc455 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.134.0.52/23"],"mac_address":"0a:58:0a:86:00:34","gateway_ips":["10.134.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.134.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.134.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.134.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.134.0.1"}],"ip_address":"10.134.0.52/23","gateway_ip":"10.134.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.134.0.52\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:86:00:34\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler-79866dc455 [e2e-llm-inference-service] uid: 5cd47741-95cf-45cb-8038-886384eb0838 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-138-159 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"5cd47741-95cf-45cb-8038-886384eb0838"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:04Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:35Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.134.0.52"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kube-api-access-h5cj8 [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type: prefix-cache-scorer\n\ [e2e-llm-inference-service] - type: max-score-picker\nschedulingProfiles:\n- name: default\n plugins:\n\ [e2e-llm-inference-service] \ - pluginRef: queue-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n\ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-h5cj8 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] fsGroup: 1000690000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa-dockercfg-pg2h6 [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:04Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:35Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:35Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] hostIP: 10.0.138.159 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.138.159 [e2e-llm-inference-service] podIP: 10.134.0.52 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.134.0.52 [e2e-llm-inference-service] startTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T13:58:04Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] imageID: ghcr.io/llm-d/llm-d-router-endpoint-picker@sha256:06b6c75d77afd0e07053402752a9736c2dfbc12a306d0d37d963aac4c1d4e6a6 [e2e-llm-inference-service] containerID: cri-o://33f394fa2f8c03e7582f20aa20c224a934ce914d80db387e9ba95f413a0f706d [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-h5cj8 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: ac9a7e46-32cf-40dc-a96f-8e59c25e1699 [e2e-llm-inference-service] resourceVersion: '71521' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa-dockercfg-pg2h6 [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa-dockercfg-pg2h6"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-x9l7t"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa-dockercfg-pg2h6 [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa-dockercfg-pg2h6 [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 42392ba0-a0c1-4bce-afa8-c06f10b9b5ad [e2e-llm-inference-service] resourceVersion: '71542' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] targetPort: grpc [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] targetPort: grpc-health [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] targetPort: metrics [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] targetPort: zmq [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] clusterIP: 172.31.209.189 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.209.189 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 7d1c3524-cc56-4901-9b40-7cb68bcab0a5 [e2e-llm-inference-service] resourceVersion: '71516' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] targetPort: 8000 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] clusterIP: 172.31.147.243 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.147.243 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 85ff9933-1007-47a5-b141-fdc5127f9b21 [e2e-llm-inference-service] resourceVersion: '71801' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:rollingUpdate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:maxSurge: {} [e2e-llm-inference-service] f:maxUnavailable: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/llm-d-inference-sim [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --model [e2e-llm-inference-service] - Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] - --mode [e2e-llm-inference-service] - random [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: RollingUpdate [e2e-llm-inference-service] rollingUpdate: [e2e-llm-inference-service] maxUnavailable: 25% [e2e-llm-inference-service] maxSurge: 25% [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff4" [e2e-llm-inference-service] has successfully progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 21aa8a13-7753-44a0-b4c1-8aebf4c701bf [e2e-llm-inference-service] resourceVersion: '72086' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n-\ [e2e-llm-inference-service] \ name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: Recreate [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler-79866dc455" [e2e-llm-inference-service] has successfully progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff4 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: cafdc8c6-2fc7-4e63-8b18-19606b53c8e4 [e2e-llm-inference-service] resourceVersion: '71799' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 797dfc7ff4 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '2' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve [e2e-llm-inference-service] uid: 85ff9933-1007-47a5-b141-fdc5127f9b21 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"85ff9933-1007-47a5-b141-fdc5127f9b21"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 797dfc7ff4 [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 797dfc7ff4 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/llm-d-inference-sim [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --model [e2e-llm-inference-service] - Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] - --mode [e2e-llm-inference-service] - random [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler-79866dc455 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 5cd47741-95cf-45cb-8038-886384eb0838 [e2e-llm-inference-service] resourceVersion: '72085' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 79866dc455 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler [e2e-llm-inference-service] uid: 21aa8a13-7753-44a0-b4c1-8aebf4c701bf [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"21aa8a13-7753-44a0-b4c1-8aebf4c701bf"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 79866dc455 [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 79866dc455 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n-\ [e2e-llm-inference-service] \ name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] serviceAccount: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-rb [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 4846fc9b-c89d-4abc-9acb-262e90041e2c [e2e-llm-inference-service] resourceVersion: '71533' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-role [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-role [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: eae06504-3c2c-4c5a-96b8-1e43a721d7b4 [e2e-llm-inference-service] resourceVersion: '71528' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] - create [e2e-llm-inference-service] - update [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - delete [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service-t5kxl [e2e-llm-inference-service] generateName: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: b241c95e-7ea4-465f-9e1d-1e015c29a9fd [e2e-llm-inference-service] resourceVersion: '72084' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T13:58:35Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] uid: 42392ba0-a0c1-4bce-afa8-c06f10b9b5ad [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:36Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"42392ba0-a0c1-4bce-afa8-c06f10b9b5ad"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.134.0.52 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l [e2e-llm-inference-service] uid: da26fdce-65db-48f8-9d58-b72f0700b8c8 [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-smngxw [e2e-llm-inference-service] generateName: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: d6d2d971-4a6a-47b5-9b43-1930475273f7 [e2e-llm-inference-service] resourceVersion: '71796' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] uid: 7d1c3524-cc56-4901-9b40-7cb68bcab0a5 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:13Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"7d1c3524-cc56-4901-9b40-7cb68bcab0a5"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.134.0.51 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z [e2e-llm-inference-service] uid: 2b688cb9-a351-4747-aa9f-01cee48d1c3d [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-rb [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 4846fc9b-c89d-4abc-9acb-262e90041e2c [e2e-llm-inference-service] resourceVersion: '71533' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:kserve-ci-e2e-test:llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-role [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-role [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: eae06504-3c2c-4c5a-96b8-1e43a721d7b4 [e2e-llm-inference-service] resourceVersion: '71528' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - create [e2e-llm-inference-service] - delete [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - update [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:58:11Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71780' [e2e-llm-inference-service] uid: 43beb71f-598d-4597-8de3-4182ca3af4be [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:04Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:58:11Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71780' [e2e-llm-inference-service] uid: 43beb71f-598d-4597-8de3-4182ca3af4be [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/Qwen/Qwen2.5-0.5B-Instruct [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:04Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:endpointPickerRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:number: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:matchLabels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPorts: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71758' [e2e-llm-inference-service] uid: fe1fbf54-8810-4730-99e4-82e0f8e31885 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] appProtocol: http [e2e-llm-inference-service] endpointPickerRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] port: [e2e-llm-inference-service] number: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPorts: [e2e-llm-inference-service] - number: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] message: Referenced by an HTTPRoute accepted by the parentRef Gateway [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] message: Referenced ExtensionRef resolved successfully [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: networking.istio.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] kind: AuthPolicy [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:04Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-policies [e2e-llm-inference-service] app.kubernetes.io/managed-by: odh-model-controller [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:rules: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:authentication: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:public: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:anonymous: {} [e2e-llm-inference-service] f:credentials: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:overrides: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fairness: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:response: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:success: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:headers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:x-gateway-inference-fairness-id: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:x-gateway-inference-objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:targetRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:04Z' [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Accepted"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Enforced"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T13:58:06Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route-authn [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71709' [e2e-llm-inference-service] uid: 9ec37320-e305-4f79-82d0-69ba6ad4b9f8 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] rules: [e2e-llm-inference-service] authentication: [e2e-llm-inference-service] public: [e2e-llm-inference-service] anonymous: {} [e2e-llm-inference-service] credentials: {} [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] overrides: [e2e-llm-inference-service] fairness: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] objective: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] response: [e2e-llm-inference-service] success: [e2e-llm-inference-service] headers: [e2e-llm-inference-service] x-gateway-inference-fairness-id: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.fairness [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] x-gateway-inference-objective: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.objective [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route [e2e-llm-inference-service] status: [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:05Z' [e2e-llm-inference-service] message: AuthPolicy has been accepted [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T13:58:06Z' [e2e-llm-inference-service] message: AuthPolicy has been successfully enforced [e2e-llm-inference-service] reason: Enforced [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Enforced [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71583' [e2e-llm-inference-service] uid: 913171f5-0635-4e30-ba86-5f11b1d2ce95 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71762' [e2e-llm-inference-service] uid: a6f4273f-14e6-4481-ac79-962f8708a8d6 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference--ip-9e4c11fc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71585' [e2e-llm-inference-service] uid: 36a426f1-b68f-468e-96b6-95b64280980c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71583' [e2e-llm-inference-service] uid: 913171f5-0635-4e30-ba86-5f11b1d2ce95 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71762' [e2e-llm-inference-service] uid: a6f4273f-14e6-4481-ac79-962f8708a8d6 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference--ip-9e4c11fc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71585' [e2e-llm-inference-service] uid: 36a426f1-b68f-468e-96b6-95b64280980c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71583' [e2e-llm-inference-service] uid: 913171f5-0635-4e30-ba86-5f11b1d2ce95 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:10Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71762' [e2e-llm-inference-service] uid: a6f4273f-14e6-4481-ac79-962f8708a8d6 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference--ip-9e4c11fc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:03Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71585' [e2e-llm-inference-service] uid: 36a426f1-b68f-468e-96b6-95b64280980c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"75012cd7-c388-45b9-bf77-6dec771cf531"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:extensionRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:portNumber: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPortNumber: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T13:58:02Z' [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-inference-pool [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] uid: 75012cd7-c388-45b9-bf77-6dec771cf531 [e2e-llm-inference-service] resourceVersion: '71555' [e2e-llm-inference-service] uid: 3d265710-73c8-4125-a3e2-f1773ab33706 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] extensionRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] portNumber: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPortNumber: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parent: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '1970-01-01T00:00:00Z' [e2e-llm-inference-service] message: Waiting for controller [e2e-llm-inference-service] reason: Pending [e2e-llm-inference-service] status: Unknown [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Status [e2e-llm-inference-service] name: default [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:43Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 797dfc7ff4 [e2e-llm-inference-service] timestamp: '2026-07-02T14:13:31Z' [e2e-llm-inference-service] window: 19.043s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 17267079n [e2e-llm-inference-service] memory: 25628Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:43Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212 [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 79866dc455 [e2e-llm-inference-service] timestamp: '2026-07-02T14:13:31Z' [e2e-llm-inference-service] window: 14.732s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 9810684n [e2e-llm-inference-service] memory: 29912Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_inference_service] [2026-07-02T14:13:43.329639] end - ❌ 943.081s: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212/v1/chat/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] _ test_llm_tls_resources[router-managed-workload-single-cpu-model-fb-opt-125m] _ [e2e-llm-inference-service] [gw1] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), chunked = False [e2e-llm-inference-service] response_conn = [e2e-llm-inference-service] preload_content = False, decode_content = False, enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] # conn.request() calls http.client.*.request, not the method in [e2e-llm-inference-service] # urllib3.request. It also calls makefile (recv) on the socket. [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn.request( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] enforce_content_length=enforce_content_length, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # We are swallowing BrokenPipeError (errno.EPIPE) since the server is [e2e-llm-inference-service] # legitimately able to close the connection after sending a valid response. [e2e-llm-inference-service] # With this behaviour, the received response is still readable. [e2e-llm-inference-service] except BrokenPipeError: [e2e-llm-inference-service] pass [e2e-llm-inference-service] except OSError as e: [e2e-llm-inference-service] # MacOS/Linux [e2e-llm-inference-service] # EPROTOTYPE and ECONNRESET are needed on macOS [e2e-llm-inference-service] # https://erickt.github.io/blog/2014/11/19/adventures-in-debugging-a-potential-osx-kernel-bug/ [e2e-llm-inference-service] # Condition changed later to emit ECONNRESET instead of only EPROTOTYPE. [e2e-llm-inference-service] if e.errno != errno.EPROTOTYPE and e.errno != errno.ECONNRESET: [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset the timeout for the recv() on the socket [e2e-llm-inference-service] read_timeout = timeout_obj.read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn.is_closed: [e2e-llm-inference-service] # In Python 3 socket.py will catch EAGAIN and return None when you [e2e-llm-inference-service] # try and read into the file pointer created by http.client, which [e2e-llm-inference-service] # instead raises a BadStatusLine exception. Instead of catching [e2e-llm-inference-service] # the exception and assuming all BadStatusLine exceptions are read [e2e-llm-inference-service] # timeouts, check for a zero timeout before making the request. [e2e-llm-inference-service] if read_timeout == 0: [e2e-llm-inference-service] raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={read_timeout})" [e2e-llm-inference-service] ) [e2e-llm-inference-service] conn.timeout = read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] # Receive the response from the server [e2e-llm-inference-service] try: [e2e-llm-inference-service] > response = conn.getresponse() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:534: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def getresponse( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] ) -> HTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get the response from the server. [e2e-llm-inference-service] [e2e-llm-inference-service] If the HTTPConnection is in the correct state, returns an instance of HTTPResponse or of whatever object is returned by the response_class variable. [e2e-llm-inference-service] [e2e-llm-inference-service] If a request has not been sent or if a previous response has not be handled, ResponseNotReady is raised. If the HTTP response indicates that the connection should be closed, then it will be closed before the response is returned. When the connection is closed, the underlying socket is closed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Raise the same error as http.client.HTTPConnection [e2e-llm-inference-service] if self._response_options is None: [e2e-llm-inference-service] raise ResponseNotReady() [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset this attribute for being used again. [e2e-llm-inference-service] resp_options = self._response_options [e2e-llm-inference-service] self._response_options = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Since the connection's timeout value may have been updated [e2e-llm-inference-service] # we need to set the timeout on the socket. [e2e-llm-inference-service] self.sock.settimeout(self.timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] # This is needed here to avoid circular import errors [e2e-llm-inference-service] from .response import HTTPResponse [e2e-llm-inference-service] [e2e-llm-inference-service] # Save a reference to the shutdown function before ownership is passed [e2e-llm-inference-service] # to httplib_response [e2e-llm-inference-service] # TODO should we implement it everywhere? [e2e-llm-inference-service] _shutdown = getattr(self.sock, "shutdown", None) [e2e-llm-inference-service] [e2e-llm-inference-service] # Get the response from http.client.HTTPConnection [e2e-llm-inference-service] > httplib_response = super().getresponse() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:571: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def getresponse(self): [e2e-llm-inference-service] """Get the response from the server. [e2e-llm-inference-service] [e2e-llm-inference-service] If the HTTPConnection is in the correct state, returns an [e2e-llm-inference-service] instance of HTTPResponse or of whatever object is returned by [e2e-llm-inference-service] the response_class variable. [e2e-llm-inference-service] [e2e-llm-inference-service] If a request has not been sent or if a previous response has [e2e-llm-inference-service] not be handled, ResponseNotReady is raised. If the HTTP [e2e-llm-inference-service] response indicates that the connection should be closed, then [e2e-llm-inference-service] it will be closed before the response is returned. When the [e2e-llm-inference-service] connection is closed, the underlying socket is closed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] # if a prior response has been completed, then forget about it. [e2e-llm-inference-service] if self.__response and self.__response.isclosed(): [e2e-llm-inference-service] self.__response = None [e2e-llm-inference-service] [e2e-llm-inference-service] # if a prior response exists, then it must be completed (otherwise, we [e2e-llm-inference-service] # cannot read this response's header to determine the connection-close [e2e-llm-inference-service] # behavior) [e2e-llm-inference-service] # [e2e-llm-inference-service] # note: if a prior response existed, but was connection-close, then the [e2e-llm-inference-service] # socket and response were made independent of this HTTPConnection [e2e-llm-inference-service] # object since a new request requires that we open a whole new [e2e-llm-inference-service] # connection [e2e-llm-inference-service] # [e2e-llm-inference-service] # this means the prior response had one of two states: [e2e-llm-inference-service] # 1) will_close: this connection was reset and the prior socket and [e2e-llm-inference-service] # response operate independently [e2e-llm-inference-service] # 2) persistent: the response was retained and we await its [e2e-llm-inference-service] # isclosed() status to become true. [e2e-llm-inference-service] # [e2e-llm-inference-service] if self.__state != _CS_REQ_SENT or self.__response: [e2e-llm-inference-service] raise ResponseNotReady(self.__state) [e2e-llm-inference-service] [e2e-llm-inference-service] if self.debuglevel > 0: [e2e-llm-inference-service] response = self.response_class(self.sock, self.debuglevel, [e2e-llm-inference-service] method=self._method) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = self.response_class(self.sock, method=self._method) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > response.begin() [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:1395: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def begin(self): [e2e-llm-inference-service] if self.headers is not None: [e2e-llm-inference-service] # we've already started reading the response [e2e-llm-inference-service] return [e2e-llm-inference-service] [e2e-llm-inference-service] # read until we get a non-100 response [e2e-llm-inference-service] while True: [e2e-llm-inference-service] > version, status, reason = self._read_status() [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:325: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def _read_status(self): [e2e-llm-inference-service] > line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1") [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:286: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] b = [e2e-llm-inference-service] [e2e-llm-inference-service] def readinto(self, b): [e2e-llm-inference-service] """Read up to len(b) bytes into the writable buffer *b* and return [e2e-llm-inference-service] the number of bytes read. If the socket is non-blocking and no bytes [e2e-llm-inference-service] are available, None is returned. [e2e-llm-inference-service] [e2e-llm-inference-service] If *b* is non-empty, a 0 return value indicates that the connection [e2e-llm-inference-service] was shutdown at the other end. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self._checkClosed() [e2e-llm-inference-service] self._checkReadable() [e2e-llm-inference-service] if self._timeout_occurred: [e2e-llm-inference-service] raise OSError("cannot read from timed out object") [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return self._sock.recv_into(b) [e2e-llm-inference-service] E TimeoutError: timed out [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/socket.py:718: TimeoutError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/tls-verification-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] > response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:787: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), chunked = False [e2e-llm-inference-service] response_conn = [e2e-llm-inference-service] preload_content = False, decode_content = False, enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] # conn.request() calls http.client.*.request, not the method in [e2e-llm-inference-service] # urllib3.request. It also calls makefile (recv) on the socket. [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn.request( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] enforce_content_length=enforce_content_length, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # We are swallowing BrokenPipeError (errno.EPIPE) since the server is [e2e-llm-inference-service] # legitimately able to close the connection after sending a valid response. [e2e-llm-inference-service] # With this behaviour, the received response is still readable. [e2e-llm-inference-service] except BrokenPipeError: [e2e-llm-inference-service] pass [e2e-llm-inference-service] except OSError as e: [e2e-llm-inference-service] # MacOS/Linux [e2e-llm-inference-service] # EPROTOTYPE and ECONNRESET are needed on macOS [e2e-llm-inference-service] # https://erickt.github.io/blog/2014/11/19/adventures-in-debugging-a-potential-osx-kernel-bug/ [e2e-llm-inference-service] # Condition changed later to emit ECONNRESET instead of only EPROTOTYPE. [e2e-llm-inference-service] if e.errno != errno.EPROTOTYPE and e.errno != errno.ECONNRESET: [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset the timeout for the recv() on the socket [e2e-llm-inference-service] read_timeout = timeout_obj.read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn.is_closed: [e2e-llm-inference-service] # In Python 3 socket.py will catch EAGAIN and return None when you [e2e-llm-inference-service] # try and read into the file pointer created by http.client, which [e2e-llm-inference-service] # instead raises a BadStatusLine exception. Instead of catching [e2e-llm-inference-service] # the exception and assuming all BadStatusLine exceptions are read [e2e-llm-inference-service] # timeouts, check for a zero timeout before making the request. [e2e-llm-inference-service] if read_timeout == 0: [e2e-llm-inference-service] raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={read_timeout})" [e2e-llm-inference-service] ) [e2e-llm-inference-service] conn.timeout = read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] # Receive the response from the server [e2e-llm-inference-service] try: [e2e-llm-inference-service] response = conn.getresponse() [e2e-llm-inference-service] except (BaseSSLError, OSError) as e: [e2e-llm-inference-service] > self._raise_timeout(err=e, url=url, timeout_value=read_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:536: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] err = TimeoutError('timed out') [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] timeout_value = 60 [e2e-llm-inference-service] [e2e-llm-inference-service] def _raise_timeout( [e2e-llm-inference-service] self, [e2e-llm-inference-service] err: BaseSSLError | OSError | SocketTimeout, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] timeout_value: _TYPE_TIMEOUT | None, [e2e-llm-inference-service] ) -> None: [e2e-llm-inference-service] """Is the error actually a timeout? Will raise a ReadTimeout or pass""" [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(err, SocketTimeout): [e2e-llm-inference-service] > raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={timeout_value})" [e2e-llm-inference-service] ) from err [e2e-llm-inference-service] E urllib3.exceptions.ReadTimeoutError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:367: ReadTimeoutError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = , stream = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), verify = '/tmp/ca.crt' [e2e-llm-inference-service] cert = None, proxies = OrderedDict() [e2e-llm-inference-service] [e2e-llm-inference-service] def send( [e2e-llm-inference-service] self, request, stream=False, timeout=None, verify=True, cert=None, proxies=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Sends PreparedRequest object. Returns Response object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param request: The :class:`PreparedRequest ` being sent. [e2e-llm-inference-service] :param stream: (optional) Whether to stream the request content. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple or urllib3 Timeout object [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether [e2e-llm-inference-service] we verify the server's TLS certificate, or a string, in which case it [e2e-llm-inference-service] must be a path to a CA bundle to use [e2e-llm-inference-service] :param cert: (optional) Any user-provided SSL certificate to be trusted. [e2e-llm-inference-service] :param proxies: (optional) The proxies dictionary to apply to the request. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn = self.get_connection_with_tls_context( [e2e-llm-inference-service] request, verify, proxies=proxies, cert=cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] except LocationValueError as e: [e2e-llm-inference-service] raise InvalidURL(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] self.cert_verify(conn, request.url, verify, cert) [e2e-llm-inference-service] url = self.request_url(request, proxies) [e2e-llm-inference-service] self.add_headers( [e2e-llm-inference-service] request, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] verify=verify, [e2e-llm-inference-service] cert=cert, [e2e-llm-inference-service] proxies=proxies, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] chunked = not (request.body is None or "Content-Length" in request.headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(timeout, tuple): [e2e-llm-inference-service] try: [e2e-llm-inference-service] connect, read = timeout [e2e-llm-inference-service] timeout = TimeoutSauce(connect=connect, read=read) [e2e-llm-inference-service] except ValueError: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, " [e2e-llm-inference-service] f"or a single float to set both timeouts to the same value." [e2e-llm-inference-service] ) [e2e-llm-inference-service] elif isinstance(timeout, TimeoutSauce): [e2e-llm-inference-service] pass [e2e-llm-inference-service] else: [e2e-llm-inference-service] timeout = TimeoutSauce(connect=timeout, read=timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > resp = conn.urlopen( [e2e-llm-inference-service] method=request.method, [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] body=request.body, [e2e-llm-inference-service] headers=request.headers, [e2e-llm-inference-service] redirect=False, [e2e-llm-inference-service] assert_same_host=False, [e2e-llm-inference-service] preload_content=False, [e2e-llm-inference-service] decode_content=False, [e2e-llm-inference-service] retries=self.max_retries, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/adapters.py:667: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=7, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/tls-verification-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=6, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/tls-verification-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=5, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/tls-verification-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=4, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/tls-verification-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=3, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/tls-verification-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=2, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/tls-verification-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=1, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/tls-verification-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/tls-verification-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/tls-verification-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] > retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:841: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] response = None [e2e-llm-inference-service] error = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] _pool = [e2e-llm-inference-service] _stacktrace = [e2e-llm-inference-service] [e2e-llm-inference-service] def increment( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str | None = None, [e2e-llm-inference-service] url: str | None = None, [e2e-llm-inference-service] response: BaseHTTPResponse | None = None, [e2e-llm-inference-service] error: Exception | None = None, [e2e-llm-inference-service] _pool: ConnectionPool | None = None, [e2e-llm-inference-service] _stacktrace: TracebackType | None = None, [e2e-llm-inference-service] ) -> Self: [e2e-llm-inference-service] """Return a new Retry object with incremented retry counters. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response: A response object, or None, if the server did not [e2e-llm-inference-service] return a response. [e2e-llm-inference-service] :type response: :class:`~urllib3.response.BaseHTTPResponse` [e2e-llm-inference-service] :param Exception error: An error encountered during the request, or [e2e-llm-inference-service] None if the response was received successfully. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: A new ``Retry`` object. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if self.total is False and error: [e2e-llm-inference-service] # Disabled, indicate to re-raise the error. [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] [e2e-llm-inference-service] total = self.total [e2e-llm-inference-service] if total is not None: [e2e-llm-inference-service] total -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] connect = self.connect [e2e-llm-inference-service] read = self.read [e2e-llm-inference-service] redirect = self.redirect [e2e-llm-inference-service] status_count = self.status [e2e-llm-inference-service] other = self.other [e2e-llm-inference-service] cause = "unknown" [e2e-llm-inference-service] status = None [e2e-llm-inference-service] redirect_location = None [e2e-llm-inference-service] [e2e-llm-inference-service] if error and self._is_connection_error(error): [e2e-llm-inference-service] # Connect retry? [e2e-llm-inference-service] if connect is False: [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif connect is not None: [e2e-llm-inference-service] connect -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error and self._is_read_error(error): [e2e-llm-inference-service] # Read retry? [e2e-llm-inference-service] if read is False or method is None or not self._is_method_retryable(method): [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif read is not None: [e2e-llm-inference-service] read -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error: [e2e-llm-inference-service] # Other retry? [e2e-llm-inference-service] if other is not None: [e2e-llm-inference-service] other -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif response and response.get_redirect_location(): [e2e-llm-inference-service] # Redirect retry? [e2e-llm-inference-service] if redirect is not None: [e2e-llm-inference-service] redirect -= 1 [e2e-llm-inference-service] cause = "too many redirects" [e2e-llm-inference-service] response_redirect_location = response.get_redirect_location() [e2e-llm-inference-service] if response_redirect_location: [e2e-llm-inference-service] redirect_location = response_redirect_location [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Incrementing because of a server error like a 500 in [e2e-llm-inference-service] # status_forcelist and the given method is in the allowed_methods [e2e-llm-inference-service] cause = ResponseError.GENERIC_ERROR [e2e-llm-inference-service] if response and response.status: [e2e-llm-inference-service] if status_count is not None: [e2e-llm-inference-service] status_count -= 1 [e2e-llm-inference-service] cause = ResponseError.SPECIFIC_ERROR.format(status_code=response.status) [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] history = self.history + ( [e2e-llm-inference-service] RequestHistory(method, url, error, status, redirect_location), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] new_retry = self.new( [e2e-llm-inference-service] total=total, [e2e-llm-inference-service] connect=connect, [e2e-llm-inference-service] read=read, [e2e-llm-inference-service] redirect=redirect, [e2e-llm-inference-service] status=status_count, [e2e-llm-inference-service] other=other, [e2e-llm-inference-service] history=history, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if new_retry.is_exhausted(): [e2e-llm-inference-service] reason = error or ResponseError(cause) [e2e-llm-inference-service] > raise MaxRetryError(_pool, url, reason) from reason # type: ignore[arg-type] [e2e-llm-inference-service] E urllib3.exceptions.MaxRetryError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/tls-verification-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/util/retry.py:519: MaxRetryError [e2e-llm-inference-service] [e2e-llm-inference-service] During handling of the above exception, another exception occurred: [e2e-llm-inference-service] [e2e-llm-inference-service] def get_successful_response(): [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_case.url_getter: [e2e-llm-inference-service] service_url = test_case.url_getter(kserve_client, test_case.llm_service) [e2e-llm-inference-service] else: [e2e-llm-inference-service] service_url = get_llm_service_url(kserve_client, test_case.llm_service) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to get service URL: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] model_url = service_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] headers = {"Content-Type": "application/json"} [e2e-llm-inference-service] if extra_headers: [e2e-llm-inference-service] headers.update(extra_headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if test_case.payload_formatter is not None: [e2e-llm-inference-service] test_payload = test_case.payload_formatter(test_case) [e2e-llm-inference-service] elif test_case.prompt is not None: [e2e-llm-inference-service] test_payload = { [e2e-llm-inference-service] "model": test_case.model_name [e2e-llm-inference-service] if not extra_headers or MODEL_ROUTING_HEADER not in extra_headers [e2e-llm-inference-service] else extra_headers[MODEL_ROUTING_HEADER], [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] else: [e2e-llm-inference-service] test_payload = None [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Calling LLM service at {model_url} with payload {test_payload}") [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_payload is not None: [e2e-llm-inference-service] > response = post_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] json_data=test_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1095: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] [e2e-llm-inference-service] def post_with_retry( [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] *, [e2e-llm-inference-service] headers: Dict = None, [e2e-llm-inference-service] json_data: Union[Dict, List] = None, [e2e-llm-inference-service] data: Union[str, bytes] = None, [e2e-llm-inference-service] stream: bool = False, [e2e-llm-inference-service] timeout: float = None, [e2e-llm-inference-service] total_retries: int = DEFAULT_RETRY_TOTAL, [e2e-llm-inference-service] backoff_factor: float = DEFAULT_RETRY_BACKOFF_FACTOR, [e2e-llm-inference-service] retry_status_codes=DEFAULT_RETRY_STATUS_CODES, [e2e-llm-inference-service] ) -> requests.Response: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Send POST request with retries for transient HTTP and network failures. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if json_data is not None and data is not None: [e2e-llm-inference-service] raise ValueError("Only one of json_data or data can be provided.") [e2e-llm-inference-service] [e2e-llm-inference-service] with _retry_session( [e2e-llm-inference-service] ["POST"], total_retries, backoff_factor, retry_status_codes [e2e-llm-inference-service] ) as session: [e2e-llm-inference-service] > return session.post( [e2e-llm-inference-service] url, [e2e-llm-inference-service] json=json_data, [e2e-llm-inference-service] data=data, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] common/http_retry.py:70: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] data = None [e2e-llm-inference-service] json = {'max_tokens': 20, 'model': 'facebook/opt-125m', 'prompt': 'KServe is a'} [e2e-llm-inference-service] kwargs = {'headers': {'Content-Type': 'application/json'}, 'stream': False, 'timeout': 60} [e2e-llm-inference-service] [e2e-llm-inference-service] def post(self, url, data=None, json=None, **kwargs): [e2e-llm-inference-service] r"""Sends a POST request. Returns :class:`Response` object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: URL for the new :class:`Request` object. [e2e-llm-inference-service] :param data: (optional) Dictionary, list of tuples, bytes, or file-like [e2e-llm-inference-service] object to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param json: (optional) json to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param \*\*kwargs: Optional arguments that ``request`` takes. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.request("POST", url, data=data, json=json, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:637: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = , method = 'POST' [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/tls-verification-test/v1/completions' [e2e-llm-inference-service] params = None, data = None, headers = {'Content-Type': 'application/json'} [e2e-llm-inference-service] cookies = None, files = None, auth = None, timeout = 60, allow_redirects = True [e2e-llm-inference-service] proxies = {}, hooks = None, stream = False, verify = None, cert = None [e2e-llm-inference-service] json = {'max_tokens': 20, 'model': 'facebook/opt-125m', 'prompt': 'KServe is a'} [e2e-llm-inference-service] [e2e-llm-inference-service] def request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] params=None, [e2e-llm-inference-service] data=None, [e2e-llm-inference-service] headers=None, [e2e-llm-inference-service] cookies=None, [e2e-llm-inference-service] files=None, [e2e-llm-inference-service] auth=None, [e2e-llm-inference-service] timeout=None, [e2e-llm-inference-service] allow_redirects=True, [e2e-llm-inference-service] proxies=None, [e2e-llm-inference-service] hooks=None, [e2e-llm-inference-service] stream=None, [e2e-llm-inference-service] verify=None, [e2e-llm-inference-service] cert=None, [e2e-llm-inference-service] json=None, [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Constructs a :class:`Request `, prepares it and sends it. [e2e-llm-inference-service] Returns :class:`Response ` object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: method for the new :class:`Request` object. [e2e-llm-inference-service] :param url: URL for the new :class:`Request` object. [e2e-llm-inference-service] :param params: (optional) Dictionary or bytes to be sent in the query [e2e-llm-inference-service] string for the :class:`Request`. [e2e-llm-inference-service] :param data: (optional) Dictionary, list of tuples, bytes, or file-like [e2e-llm-inference-service] object to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param json: (optional) json to send in the body of the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param headers: (optional) Dictionary of HTTP Headers to send with the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param cookies: (optional) Dict or CookieJar object to send with the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param files: (optional) Dictionary of ``'filename': file-like-objects`` [e2e-llm-inference-service] for multipart encoding upload. [e2e-llm-inference-service] :param auth: (optional) Auth tuple or callable to enable [e2e-llm-inference-service] Basic/Digest/Custom HTTP Auth. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple [e2e-llm-inference-service] :param allow_redirects: (optional) Set to True by default. [e2e-llm-inference-service] :type allow_redirects: bool [e2e-llm-inference-service] :param proxies: (optional) Dictionary mapping protocol or protocol and [e2e-llm-inference-service] hostname to the URL of the proxy. [e2e-llm-inference-service] :param hooks: (optional) Dictionary mapping hook name to one event or [e2e-llm-inference-service] list of events, event must be callable. [e2e-llm-inference-service] :param stream: (optional) whether to immediately download the response [e2e-llm-inference-service] content. Defaults to ``False``. [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether we verify [e2e-llm-inference-service] the server's TLS certificate, or a string, in which case it must be a path [e2e-llm-inference-service] to a CA bundle to use. Defaults to ``True``. When set to [e2e-llm-inference-service] ``False``, requests will accept any TLS certificate presented by [e2e-llm-inference-service] the server, and will ignore hostname mismatches and/or expired [e2e-llm-inference-service] certificates, which will make your application vulnerable to [e2e-llm-inference-service] man-in-the-middle (MitM) attacks. Setting verify to ``False`` [e2e-llm-inference-service] may be useful during local development or testing. [e2e-llm-inference-service] :param cert: (optional) if String, path to ssl client cert file (.pem). [e2e-llm-inference-service] If Tuple, ('cert', 'key') pair. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Create the Request. [e2e-llm-inference-service] req = Request( [e2e-llm-inference-service] method=method.upper(), [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] files=files, [e2e-llm-inference-service] data=data or {}, [e2e-llm-inference-service] json=json, [e2e-llm-inference-service] params=params or {}, [e2e-llm-inference-service] auth=auth, [e2e-llm-inference-service] cookies=cookies, [e2e-llm-inference-service] hooks=hooks, [e2e-llm-inference-service] ) [e2e-llm-inference-service] prep = self.prepare_request(req) [e2e-llm-inference-service] [e2e-llm-inference-service] proxies = proxies or {} [e2e-llm-inference-service] [e2e-llm-inference-service] settings = self.merge_environment_settings( [e2e-llm-inference-service] prep.url, proxies, stream, verify, cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Send the request. [e2e-llm-inference-service] send_kwargs = { [e2e-llm-inference-service] "timeout": timeout, [e2e-llm-inference-service] "allow_redirects": allow_redirects, [e2e-llm-inference-service] } [e2e-llm-inference-service] send_kwargs.update(settings) [e2e-llm-inference-service] > resp = self.send(prep, **send_kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:589: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = [e2e-llm-inference-service] kwargs = {'cert': None, 'proxies': OrderedDict(), 'stream': False, 'timeout': 60, ...} [e2e-llm-inference-service] allow_redirects = True, stream = False, hooks = {'response': []} [e2e-llm-inference-service] adapter = [e2e-llm-inference-service] start = 1783001607.5304596 [e2e-llm-inference-service] [e2e-llm-inference-service] def send(self, request, **kwargs): [e2e-llm-inference-service] """Send a given PreparedRequest. [e2e-llm-inference-service] [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Set defaults that the hooks can utilize to ensure they always have [e2e-llm-inference-service] # the correct parameters to reproduce the previous request. [e2e-llm-inference-service] kwargs.setdefault("stream", self.stream) [e2e-llm-inference-service] kwargs.setdefault("verify", self.verify) [e2e-llm-inference-service] kwargs.setdefault("cert", self.cert) [e2e-llm-inference-service] if "proxies" not in kwargs: [e2e-llm-inference-service] kwargs["proxies"] = resolve_proxies(request, self.proxies, self.trust_env) [e2e-llm-inference-service] [e2e-llm-inference-service] # It's possible that users might accidentally send a Request object. [e2e-llm-inference-service] # Guard against that specific failure case. [e2e-llm-inference-service] if isinstance(request, Request): [e2e-llm-inference-service] raise ValueError("You can only send PreparedRequests.") [e2e-llm-inference-service] [e2e-llm-inference-service] # Set up variables needed for resolve_redirects and dispatching of hooks [e2e-llm-inference-service] allow_redirects = kwargs.pop("allow_redirects", True) [e2e-llm-inference-service] stream = kwargs.get("stream") [e2e-llm-inference-service] hooks = request.hooks [e2e-llm-inference-service] [e2e-llm-inference-service] # Get the appropriate adapter to use [e2e-llm-inference-service] adapter = self.get_adapter(url=request.url) [e2e-llm-inference-service] [e2e-llm-inference-service] # Start time (approximately) of the request [e2e-llm-inference-service] start = preferred_clock() [e2e-llm-inference-service] [e2e-llm-inference-service] # Send the request [e2e-llm-inference-service] > r = adapter.send(request, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:703: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = , stream = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), verify = '/tmp/ca.crt' [e2e-llm-inference-service] cert = None, proxies = OrderedDict() [e2e-llm-inference-service] [e2e-llm-inference-service] def send( [e2e-llm-inference-service] self, request, stream=False, timeout=None, verify=True, cert=None, proxies=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Sends PreparedRequest object. Returns Response object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param request: The :class:`PreparedRequest ` being sent. [e2e-llm-inference-service] :param stream: (optional) Whether to stream the request content. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple or urllib3 Timeout object [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether [e2e-llm-inference-service] we verify the server's TLS certificate, or a string, in which case it [e2e-llm-inference-service] must be a path to a CA bundle to use [e2e-llm-inference-service] :param cert: (optional) Any user-provided SSL certificate to be trusted. [e2e-llm-inference-service] :param proxies: (optional) The proxies dictionary to apply to the request. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn = self.get_connection_with_tls_context( [e2e-llm-inference-service] request, verify, proxies=proxies, cert=cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] except LocationValueError as e: [e2e-llm-inference-service] raise InvalidURL(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] self.cert_verify(conn, request.url, verify, cert) [e2e-llm-inference-service] url = self.request_url(request, proxies) [e2e-llm-inference-service] self.add_headers( [e2e-llm-inference-service] request, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] verify=verify, [e2e-llm-inference-service] cert=cert, [e2e-llm-inference-service] proxies=proxies, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] chunked = not (request.body is None or "Content-Length" in request.headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(timeout, tuple): [e2e-llm-inference-service] try: [e2e-llm-inference-service] connect, read = timeout [e2e-llm-inference-service] timeout = TimeoutSauce(connect=connect, read=read) [e2e-llm-inference-service] except ValueError: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, " [e2e-llm-inference-service] f"or a single float to set both timeouts to the same value." [e2e-llm-inference-service] ) [e2e-llm-inference-service] elif isinstance(timeout, TimeoutSauce): [e2e-llm-inference-service] pass [e2e-llm-inference-service] else: [e2e-llm-inference-service] timeout = TimeoutSauce(connect=timeout, read=timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] resp = conn.urlopen( [e2e-llm-inference-service] method=request.method, [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] body=request.body, [e2e-llm-inference-service] headers=request.headers, [e2e-llm-inference-service] redirect=False, [e2e-llm-inference-service] assert_same_host=False, [e2e-llm-inference-service] preload_content=False, [e2e-llm-inference-service] decode_content=False, [e2e-llm-inference-service] retries=self.max_retries, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] except (ProtocolError, OSError) as err: [e2e-llm-inference-service] raise ConnectionError(err, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] except MaxRetryError as e: [e2e-llm-inference-service] if isinstance(e.reason, ConnectTimeoutError): [e2e-llm-inference-service] # TODO: Remove this in 3.0.0: see #2811 [e2e-llm-inference-service] if not isinstance(e.reason, NewConnectionError): [e2e-llm-inference-service] raise ConnectTimeout(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, ResponseError): [e2e-llm-inference-service] raise RetryError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, _ProxyError): [e2e-llm-inference-service] raise ProxyError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, _SSLError): [e2e-llm-inference-service] # This branch is for urllib3 v1.22 and later. [e2e-llm-inference-service] raise SSLError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] > raise ConnectionError(e, request=request) [e2e-llm-inference-service] E requests.exceptions.ConnectionError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/tls-verification-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/adapters.py:700: ConnectionError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name=... {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] service_name="tls-verification-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_tls_resources(test_case: TestCase): [e2e-llm-inference-service] """Verify that TLS-related resources (DestinationRules, cert secrets, service port) [e2e-llm-inference-service] are correctly present or absent based on the enableLLMInferenceServiceTLS flag.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] tls_enabled = _is_tls_enabled() [e2e-llm-inference-service] logger.info(f"enableLLMInferenceServiceTLS = {tls_enabled}") [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] if not test_case.llm_service.metadata.annotations: [e2e-llm-inference-service] test_case.llm_service.metadata.annotations = {} [e2e-llm-inference-service] [e2e-llm-inference-service] test_case.llm_service.metadata.annotations[ [e2e-llm-inference-service] "security.opendatahub.io/enable-auth" [e2e-llm-inference-service] ] = "false" [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] > wait_for_model_response(kserve_client, test_case, test_case.wait_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_tls.py:143: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, TestCase(base_refs=['router-managed', 'workload-sin... {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m'), 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_model_response' [e2e-llm-inference-service] timestamp_start = '2026-07-02T14:13:27.519014', start_time = 1783001607.5192838 [e2e-llm-inference-service] duration = 904.6867213249207, timestamp_end = '2026-07-02T14:28:32.206009' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name=... {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] timeout_seconds = 900, extra_headers = None [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_model_response( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] test_case: TestCase, # noqa: F811 [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] extra_headers: Optional[Dict[str, str]] = None, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def get_successful_response(): [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_case.url_getter: [e2e-llm-inference-service] service_url = test_case.url_getter(kserve_client, test_case.llm_service) [e2e-llm-inference-service] else: [e2e-llm-inference-service] service_url = get_llm_service_url(kserve_client, test_case.llm_service) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to get service URL: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] model_url = service_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] headers = {"Content-Type": "application/json"} [e2e-llm-inference-service] if extra_headers: [e2e-llm-inference-service] headers.update(extra_headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if test_case.payload_formatter is not None: [e2e-llm-inference-service] test_payload = test_case.payload_formatter(test_case) [e2e-llm-inference-service] elif test_case.prompt is not None: [e2e-llm-inference-service] test_payload = { [e2e-llm-inference-service] "model": test_case.model_name [e2e-llm-inference-service] if not extra_headers or MODEL_ROUTING_HEADER not in extra_headers [e2e-llm-inference-service] else extra_headers[MODEL_ROUTING_HEADER], [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] else: [e2e-llm-inference-service] test_payload = None [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Calling LLM service at {model_url} with payload {test_payload}") [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_payload is not None: [e2e-llm-inference-service] response = post_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] json_data=test_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = get_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] logger.error(f"❌ Failed to call model: {e}") [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to call model: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Model response is {response.status_code}: {response.text[:500]}") [e2e-llm-inference-service] [e2e-llm-inference-service] if 200 <= response.status_code < 300: [e2e-llm-inference-service] return response [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Service returned {response.status_code}: {response.text}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] > response = wait_for(get_successful_response, timeout=timeout_seconds, interval=5.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1119: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .get_successful_response at 0x7f3d99c50c20> [e2e-llm-inference-service] timeout = 900, interval = 5.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1215: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def get_successful_response(): [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_case.url_getter: [e2e-llm-inference-service] service_url = test_case.url_getter(kserve_client, test_case.llm_service) [e2e-llm-inference-service] else: [e2e-llm-inference-service] service_url = get_llm_service_url(kserve_client, test_case.llm_service) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to get service URL: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] model_url = service_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] headers = {"Content-Type": "application/json"} [e2e-llm-inference-service] if extra_headers: [e2e-llm-inference-service] headers.update(extra_headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if test_case.payload_formatter is not None: [e2e-llm-inference-service] test_payload = test_case.payload_formatter(test_case) [e2e-llm-inference-service] elif test_case.prompt is not None: [e2e-llm-inference-service] test_payload = { [e2e-llm-inference-service] "model": test_case.model_name [e2e-llm-inference-service] if not extra_headers or MODEL_ROUTING_HEADER not in extra_headers [e2e-llm-inference-service] else extra_headers[MODEL_ROUTING_HEADER], [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] else: [e2e-llm-inference-service] test_payload = None [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Calling LLM service at {model_url} with payload {test_payload}") [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_payload is not None: [e2e-llm-inference-service] response = post_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] json_data=test_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = get_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] logger.error(f"❌ Failed to call model: {e}") [e2e-llm-inference-service] > raise AssertionError(f"❌ Failed to call model: {e}") from e [e2e-llm-inference-service] E AssertionError: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/tls-verification-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1109: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig router-managed-tls-verification-86160844 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig router-managed-tls-verification-86160844 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig router-managed-tls-verification-86160844 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig workload-single-cpu-tls-verific-00217f09 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig workload-single-cpu-tls-verific-00217f09 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig workload-single-cpu-tls-verific-00217f09 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig model-fb-opt-125m-tls-verificat-7c5fca28 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig model-fb-opt-125m-tls-verificat-7c5fca28 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig model-fb-opt-125m-tls-verificat-7c5fca28 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_tls_resources] [2026-07-02T14:11:22.162417] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name='tls-verification-test', endpoint='/v1/completions', max_tokens=20, payload_formatter=, response_assertion=.response_assertion at 0x7f3d9afec4a0>, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': None, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'tls-verification-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-tls-verification-86160844'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.test_llm_tls:test_llm_tls.py:123 enableLLMInferenceServiceTLS = True [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-02T14:11:22.214407] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'tls-verification-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-tls-verification-86160844'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-02T14:11:22.312487] end - ✅ in 0.098s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-02T14:11:22.312722] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'tls-verification-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-tls-verification-86160844'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-02T14:11:28Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/tls-verification-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T14:11:28Z', 'message': 'Inference Pool kserve-ci-e2e-test/tls-verification-test-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T14:11:28Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:11:28Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T14:11:28Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/tls-verification-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T14:11:28Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/tls-verification-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T14:11:28Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:11:28Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'Ready', 'RouterReady', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-02T14:11:53Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T14:11:53Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T14:11:53Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:11:28Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T14:11:53Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T14:11:53Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T14:11:53Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:11:53Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'Ready', 'WorkloadsReady'}, expected {'Ready', 'RouterReady', 'WorkloadsReady'}, got [{'lastTransitionTime': '2026-07-02T14:11:53Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T14:11:53Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T14:11:53Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:11:28Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T14:11:53Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T14:12:00Z', 'status': 'True', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T14:12:00Z', 'severity': 'Info', 'status': 'True', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:11:53Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [wait_for_llm_isvc_ready] [2026-07-02T14:13:27.518875] end - ✅ in 125.206s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_model_response] [2026-07-02T14:13:27.519014] start - args=(, TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name='tls-verification-test', endpoint='/v1/completions', max_tokens=20, payload_formatter=, response_assertion=.response_assertion at 0x7f3d9afec4a0>, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'tls-verification-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-tls-verification-86160844'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m'), 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [get_llm_service_url] [2026-07-02T14:13:27.519287] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'tls-verification-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-tls-verification-86160844'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [get_llm_service_url] [2026-07-02T14:13:27.528609] end - ✅ in 0.009s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1092 Calling LLM service at http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/tls-verification-test/v1/completions with payload {'model': 'facebook/opt-125m', 'prompt': 'KServe is a', 'max_tokens': 20} [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=7, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/tls-verification-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=6, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/tls-verification-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=5, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/tls-verification-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=4, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/tls-verification-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=3, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/tls-verification-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=2, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/tls-verification-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=1, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/tls-verification-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=0, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/tls-verification-test/v1/completions [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1108 ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/tls-verification-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1219 Timed out waiting: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/tls-verification-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_model_response] [2026-07-02T14:28:32.206009] end - ❌ 904.687s: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/tls-verification-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] ERROR e2e.llmisvc.test_llm_tls:test_llm_tls.py:148 Failed TLS verification for tls-verification-test: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/tls-verification-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1240 🔍 # Diagnostics for 'tls-verification-test' in 'kserve-ci-e2e-test' [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1241 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1242 # LLMInferenceService tls-verification-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1245 apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] security.opendatahub.io/enable-auth: 'false' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:22Z' [e2e-llm-inference-service] finalizers: [e2e-llm-inference-service] - serving.kserve.io/llmisvc-finalizer [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:security.opendatahub.io/enable-auth: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:baseRefs: {} [e2e-llm-inference-service] manager: OpenAPI-Generator [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:22Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:finalizers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] v:"serving.kserve.io/llmisvc-finalizer": {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:22Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:addresses: {} [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-router-route: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-scheduler: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-tracing: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-worker-data-parallel: {} [e2e-llm-inference-service] f:appliedConfigs: {} [e2e-llm-inference-service] f:conditions: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:router: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:gateways: {} [e2e-llm-inference-service] f:scheduler: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:inferencePool: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:service: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:url: {} [e2e-llm-inference-service] f:workloads: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:primary: {} [e2e-llm-inference-service] f:scheduler: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:13:27Z' [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] resourceVersion: '83815' [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] spec: [e2e-llm-inference-service] baseRefs: [e2e-llm-inference-service] - name: router-managed-tls-verification-86160844 [e2e-llm-inference-service] - name: workload-single-cpu-tls-verific-00217f09 [e2e-llm-inference-service] - name: model-fb-opt-125m-tls-verificat-7c5fca28 [e2e-llm-inference-service] model: [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uri: '' [e2e-llm-inference-service] status: [e2e-llm-inference-service] addresses: [e2e-llm-inference-service] - name: gateway-external-model-routing [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/ [e2e-llm-inference-service] - name: gateway-external [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/tls-verification-test [e2e-llm-inference-service] - name: gateway-internal-model-routing [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/ [e2e-llm-inference-service] - name: gateway-internal [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/kserve-ci-e2e-test/tls-verification-test [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-template: kserve-config-llm-decode-template [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-worker-data-parallel: kserve-config-llm-decode-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-template: kserve-config-llm-prefill-template [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-worker-data-parallel: kserve-config-llm-prefill-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-router-route: kserve-config-llm-router-route [e2e-llm-inference-service] serving.kserve.io/config-llm-scheduler: kserve-config-llm-scheduler [e2e-llm-inference-service] serving.kserve.io/config-llm-template: kserve-config-llm-template [e2e-llm-inference-service] serving.kserve.io/config-llm-tracing: kserve-config-llm-tracing [e2e-llm-inference-service] serving.kserve.io/config-llm-worker-data-parallel: kserve-config-llm-worker-data-parallel [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:53Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: HTTPRoutesReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:53Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: InferencePoolReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:13:27Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: MainWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:28Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: PresetsCombined [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:13:27Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Ready [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: RouterReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: SchedulerWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:13:27Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: WorkloadsReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/tls-verification-test [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:44 TIME NAMESPACE SOURCE TYPE REASON MESSAGE [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:45 -------------------------------------------------------------------------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-6599bc8fd-8sr69 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-6599bc8fd-8sr69 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-router-scheduler-6fcbb55f87 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-6599bc8fd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:52 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy auth-disabled-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "auth-disabled-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-disabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-disabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-disabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-disabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:56 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-disabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-67c57d579b-gtwnj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:05 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 27.649s (27.649s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.35:8000/health": dial tcp 10.132.0.35:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-67c57d579b-gtwnj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.31/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.181s (1.181s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-router-scheduler-688c79755d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-67c57d579b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-enabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-enabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-enabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-enabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-enabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-enabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-6cd6c5884-rrbrr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-6cd6c5884-rrbrr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-router-scheduler-5b958c76d8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-6cd6c5884 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-invalid-token-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-invalid-token-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-invalid-token-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-invalid-token-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-invalid-token-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:24 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-invalid-token-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:25 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.52:8001/health": dial tcp 10.132.0.52:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.53/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.53:8000/health": dial tcp 10.132.0.53:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-router-scheduler-f7968dfdd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-79d6b9cc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:31 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-pd-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-pd-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:43 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-98989556b-vl89r to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:14:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.45:8000/health": dial tcp 10.132.0.45:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-98989556b-vl89r [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-router-scheduler-579b7545d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-98989556b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:57 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:15:01 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/e2e-pvc-model-download-4s522 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:26 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test job-controller Normal SuccessfulCreate Created pod: e2e-pvc-model-download-4s522 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:36 kserve-ci-e2e-test job-controller Normal Completed Job completed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:20 kserve-ci-e2e-test persistentvolume-controller Normal WaitForFirstConsumer waiting for first consumer to be created before binding [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test persistentvolume-controller Normal ExternalProvisioning Waiting for a volume to be created either by the external provisioner 'ebs.csi.aws.com' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal Provisioning External provisioner is provisioning volume for claim "kserve-ci-e2e-test/e2e-pvc-model-storage" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal ProvisioningSucceeded Successfully provisioned volume pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.094s (1.094s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:28 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec0c69dceeb48768325d1a53a749e65786-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.30/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedMount MountVolume.SetUp failed for volume "tls-certs" : secret "gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.60/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:46 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.60:8000/health": dial tcp 10.132.0.60:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.60:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:32 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-3c960099-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-3c960099-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:02 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-3c960099] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" in 808ms (808ms including waiting). Image size: 44914394 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.48:8001/health": dial tcp 10.132.0.48:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.49:8000/health": dial tcp 10.132.0.49:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5b86594bc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:44 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:52 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-50bc673d] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:07:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.44:8000/health": dial tcp 10.132.0.44:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:07:54 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-87882a8e] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 44.193s (44.193s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.62/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:40:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.62:8000/health": dial tcp 10.132.0.62:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 7f6c8cfc6 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:40:17 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test386bd5808c5c450f11fd8632e1f821fb-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:58 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-dc21cb14] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:20 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.45:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:27 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-e95b1dc1] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.42:8000/health": dial tcp 10.132.0.42:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler-67f9b884d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:01 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:19 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-7ca60146] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:44 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler-5444b4dbf from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:57 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:55 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-ba4d693a] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.55/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.54/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:25 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 686d468674 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:35 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-2577e794-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:53 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-2577e794] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:56 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.47:8000/health": dial tcp 10.132.0.47:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:51 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-59b9d263-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-59b9d263-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:25 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-59b9d263] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.50:8001/health": dial tcp 10.132.0.50:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.51:8000/health": dial tcp 10.132.0.51:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-79497db4cc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:56 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-e8706282-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:09 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-e8706282] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test80f99418f0f3eb4c80f02e6144a6d4f2-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:36 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4c1ba212] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:30 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.34:8000/health": dial tcp 10.133.0.34:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:36 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:14 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4f8c0978] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.32/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:50 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:48 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-a50492e9] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:40 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-4b931143-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-4b931143-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:21 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-4b931143] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:25 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.42:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-5b1e8f15-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-5b1e8f15-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:22 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-5b1e8f15] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-e45d1f79-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-e45d1f79-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:20 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-e45d1f79] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler-6fcb489785 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.566s (1.566s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-7dbcb75dbc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0, llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler-79866dc455 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:03 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedToRetrieveImagePullSecret Unable to retrieve some image pull secrets (llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa-dockercfg-pg2h6); attempting to pull the image may not succeed. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc44d181485fad85e662eb092f3749502f-kserve-router-scheduler-69bf65579d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler-5dd88bfbb7 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler-5f555d4d85 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler-54ccf64f4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler-6d86bd4d9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-67b4bb9646 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9, llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler-68cc9685d6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler-749449dbc8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-ccbb969bd-65ln7 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.63/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:56:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.63:8000/health": dial tcp 10.132.0.63:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-multiple-adapters-test-kserve-ccbb969bd-65ln7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-multiple-adapters-test-kserve-ccbb969bd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:08 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-multiple-adapters-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-multiple-adapters-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-multiple-adapters-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:56:42 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-multiple-adapters-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.61/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-single-adapter-test-kserve-85cd6c88dc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-single-adapter-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-single-adapter-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-single-adapter-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-single-adapter-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.29/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.483s (1.483s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.29:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 3.651s (3.651s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.88s (1.88s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 2.962s (2.962s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.723s (1.723s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" in 30.447s (30.447s including waiting). Image size: 2989890188 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Liveness probe failed: timeout: failed to connect service "10.132.0.34:9003" within 1s: context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-router-scheduler-77d88bdcf4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-7c94f9d7c6 from 0 to 2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy precise-prefix-cache-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "precise-prefix-cache-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/precise-prefix-cache-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/precise-prefix-cache-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/precise-prefix-cache-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/precise-prefix-cache-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:15 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [precise-prefix-cache-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:16 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-1-openshift-default-799f46c59b-fnbwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.28/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:31 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "http://10.133.0.28:15021/healthz/ready": context deadline exceeded (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning BackOff Back-off restarting failed container istio-proxy in pod router-gateway-1-openshift-default-799f46c59b-fnbwc_kserve-ci-e2e-test(f244ca92-1cf1-429e-886d-e0d585d0175f) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.28:15021/healthz/ready": dial tcp 10.133.0.28:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-1-openshift-default-799f46c59b-fnbwc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:58 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-1-openshift-default-799f46c59b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:37 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-2-openshift-default-54c789bdc6-w57g9 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:19 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.44:15021/healthz/ready": dial tcp 10.133.0.44:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-2-openshift-default-54c789bdc6-w57g9 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-2-openshift-default-54c789bdc6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:17 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.56:8001/health": dial tcp 10.132.0.56:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.57/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.57:8000/health": dial tcp 10.132.0.57:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-prefill-778968b98b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-router-scheduler-5c684fd54d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-694c6ddf4c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:33 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:47 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-587fcd8566-2pqvl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:20:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:29 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-587fcd8566-2pqvl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-router-scheduler-74545d8489 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-587fcd8566 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:05 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.66/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-configmap-ref-test-kserve-router-scheduler-68f876cc5d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-configmap-ref-test-kserve-744bf59fbd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy scheduler-configmap-ref-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "scheduler-configmap-ref-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-configmap-ref-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [scheduler-configmap-ref-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-7c6c58cc96-l62qh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-7c6c58cc96-l62qh [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-router-scheduler-d9ddf6c4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-7c6c58cc96 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy scheduler-inline-config-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "scheduler-inline-config-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/scheduler-inline-config-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/scheduler-inline-config-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/scheduler-inline-config-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/scheduler-inline-config-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [scheduler-inline-config-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-6hzlt to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.59/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:51 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-f2dmd to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.58/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-f2dmd [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-6hzlt [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:07 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy stop-feature-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "stop-feature-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-stop-feature-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/stop-feature-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/stop-feature-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/stop-feature-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [stop-feature-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.InferencePool kserve-ci-e2e-test/stop-feature-test-inference-pool [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:56 kserve-ci-e2e-test LLMInferenceServiceController Warning LLMInferenceServiceNotReady LLMInferenceService [stop-feature-test] is no longer Ready because of: GatewaysReady, HTTPRoutesReady, InferencePoolReady, MainWorkloadReady, PrefillWorkerWorkloadReady, PrefillWorkloadReady, RouterReady, SchedulerWorkloadReady, WorkerWorkloadReady, WorkloadsReady [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/tls-verification-test-kserve-6d6946dcd6-bdddw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.64/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.64:8000/health": dial tcp 10.132.0.64:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: tls-verification-test-kserve-6d6946dcd6-bdddw [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.65/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set tls-verification-test-kserve-router-scheduler-7f955879f5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set tls-verification-test-kserve-6d6946dcd6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:24 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy tls-verification-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/tls-verification-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "tls-verification-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/tls-verification-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/tls-verification-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-tls-verification-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/tls-verification-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/tls-verification-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/tls-verification-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/tls-verification-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/tls-verification-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [tls-verification-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod tls-verification-test-kserve-6d6946dcd6-bdddw (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### init-container 'storage-initializer' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 2026-07-02 14:11:27.084 1 storage.initializer INFO [initializer-entrypoint:():17] Initializing, args: (src_uri, dest_path): [('hf://facebook/opt-125m', '/mnt/models')] [e2e-llm-inference-service] 2026-07-02 14:11:27.084 1 storage.initializer INFO [kserve_storage.py:download():166] Copying contents of hf://facebook/opt-125m to local [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/wPaCkH-WbT7GsmxMKKrNZTV4nSM=.ac481c8eb05e4d2496fbe076a38a7b4835dd733d.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_e232e326-19a5-4281-be68-4519fbcba120'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/5HHJ6px3_ZRDOG3OxNZMhuycwOk=.a591333512516f58bf2002045dece909a0ccdb8b.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_e891f642-d9ec-40e3-8687-d4a72c7ab574'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/Xn7B-BWUGOee2Y6hCZtEhtFu4BE=.38c05904caf6e5b9f04ecda5c973d77e6c1da151.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_e1c1851a-0ea3-4094-8df1-f90f15a7e421'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/8_PA_wEVGiVa2goH2H4KQOQpvVY=.b3fb716a3024261980becb2382e31a3780985130.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_cee6f5d5-02f1-43e2-ba0b-d3eefca90a4f'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/gPcsVCQDYDHk-_n0G9uADl7PXIM=.61c60ec52ed43038fff0fbbd68b080c94b0d94b4c8458dbd65965f9b17631c89.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_2f508c89-01c0-4c6d-bc89-713bc04161fc'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/3EVKVggOldJcKSsGjSdoUCN1AyQ=.cf739e3ba86db7791ebab2828cc34b8a5acd3a86.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_174e642d-2d0f-455a-84a1-6962ac4a7339'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/PtHk0z_I45atnj23IIRhTExwT3w=.226b0752cac7789c48f0cb3ec53eda48b7be36cc.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_6298e155-aacc-4cd6-bace-d2b51a07dc89'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/Q1p2l2BzM1m6P5jKvr8WTq1TUio=.2d74da6615135c58cf3cf9ad4cb11e7c613ff9e55fe658a47ab83b6c8d1174a9.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_ac8511a5-d43c-4cbc-9a64-534f27f20d49'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/ahkChHUJFxEmOdq5GDFEmerRzCY=.5dfa36546b8eddce0e04df3133c30df43fcc3828.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_e3a0dfc5-6d62-4c54-ae7c-93274e193f39'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/a7eHxRFT3OeMBIFg52k2nfj5m7w=.db7090b0c8b34dd957a7e0656c718f978f9203cc874018f37dda44108be5970a.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_23acaf52-d8cf-4e89-94c2-ef07cd14bd40'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/vzaExXFZNBay89bvlQv-ZcI6BTg=.27c24ca9d908d0b678b20c698aeb9e950c44d865.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_0b5af66c-5667-4096-a081-60ae7b204b57'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/j3m-Hy6QvBddw8RXA1uSWl1AJ0c=.0a39732b2d8be8e493cab3da68b68cc3e28221de.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_1c4a5f1f-6505-483f-a695-62fa096f4626'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] 2026-07-02 14:11:30.909 1 storage.initializer INFO [kserve_storage.py:download():234] Successfully copied hf://facebook/opt-125m to /mnt/models [e2e-llm-inference-service] 2026-07-02 14:11:30.909 1 storage.initializer INFO [kserve_storage.py:download():235] Model downloaded in 3.8248580700001185 seconds. [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 (APIServer pid=1) DEBUG 07-02 14:25:46 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:47 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:48 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:49 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:50 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:51 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:51 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:52 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:53 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:54 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:55 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:56 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:56 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:57 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:58 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:25:59 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:00 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:01 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:01 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:02 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:03 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:04 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:05 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:06 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:06 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:07 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:08 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:09 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:10 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:11 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:11 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:12 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:13 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:14 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:15 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:16 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:16 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:17 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:18 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:19 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:20 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:21 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:21 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:22 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:23 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:24 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:25 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:26 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:26 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:27 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:28 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:29 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:30 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:31 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:31 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:32 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:33 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:34 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:35 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:36 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:36 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:37 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:38 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:39 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:40 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:41 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:41 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:42 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:43 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:44 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:45 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:46 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:46 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:47 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:48 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:49 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:50 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:51 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:51 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:52 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:53 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:54 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:55 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:56 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:56 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:57 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:58 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:26:59 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:00 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:01 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:01 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:02 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:03 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:04 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:05 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:06 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:06 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:07 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:08 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:09 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:10 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:11 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:11 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:12 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:13 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:14 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:15 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:16 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:16 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:17 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:18 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:19 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:20 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:21 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:21 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:22 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:23 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:24 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:25 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:26 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:26 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:27 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:28 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:29 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:30 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:31 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:31 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:32 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:33 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:34 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:35 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:36 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:36 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:37 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:38 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:39 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:40 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:41 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:41 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:42 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:43 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:44 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:45 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:46 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:46 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:47 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:48 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:49 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:50 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:51 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:51 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:52 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:53 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:54 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:55 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:56 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:56 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:57 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:58 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:27:59 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:00 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:01 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:01 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:02 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:03 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:04 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:05 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:06 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:06 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:07 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:08 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:09 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:10 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:11 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:11 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:12 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:13 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:14 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:15 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:16 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:16 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:17 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:18 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:19 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:20 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:21 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:21 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:22 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:23 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:24 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:25 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:26 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:26 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:27 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:28 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:29 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:30 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:31 [v1/metrics/loggers.py:259] Engine 000: Avg prompt throughput: 0.0 tokens/s, Avg generation throughput: 0.0 tokens/s, Running: 0 reqs, Waiting: 0 reqs, GPU KV cache usage: 0.0%, Prefix cache hit rate: 0.0% [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:31 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] (APIServer pid=1) DEBUG 07-02 14:28:32 [v1/engine/async_llm.py:875] Called check_health. [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### init-container 'storage-initializer' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 2026-07-02 14:11:27.419 1 storage.initializer INFO [initializer-entrypoint:():17] Initializing, args: (src_uri, dest_path): [('hf://facebook/opt-125m', '/mnt/models')] [e2e-llm-inference-service] 2026-07-02 14:11:27.419 1 storage.initializer INFO [kserve_storage.py:download():166] Copying contents of hf://facebook/opt-125m to local [e2e-llm-inference-service] 2026-07-02 14:11:27.419 1 storage.initializer INFO [kserve_storage.py:download():169] Allow patterns: ['tokenizer.json', 'tokenizer_config.json', 'special_tokens_map.json', 'vocab.json', 'merges.txt', 'config.json', 'generation_config.json'] [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/8_PA_wEVGiVa2goH2H4KQOQpvVY=.b3fb716a3024261980becb2382e31a3780985130.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_715a38dd-c230-49af-89f0-b6e8c7e2d844'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/3EVKVggOldJcKSsGjSdoUCN1AyQ=.cf739e3ba86db7791ebab2828cc34b8a5acd3a86.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_1b5d5dcc-b359-4e66-ac2b-10412407c52a'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/PtHk0z_I45atnj23IIRhTExwT3w=.226b0752cac7789c48f0cb3ec53eda48b7be36cc.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_2ecb56c7-72ea-47f3-a5b1-d00d0190323c'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/ahkChHUJFxEmOdq5GDFEmerRzCY=.5dfa36546b8eddce0e04df3133c30df43fcc3828.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_be04708b-7e97-4682-b04d-2f5b3903e197'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/vzaExXFZNBay89bvlQv-ZcI6BTg=.27c24ca9d908d0b678b20c698aeb9e950c44d865.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_ab9e0fbb-08dc-4735-9c4e-ddeaea96d367'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] Could not set the permissions on the file '/mnt/models/.cache/huggingface/download/j3m-Hy6QvBddw8RXA1uSWl1AJ0c=.0a39732b2d8be8e493cab3da68b68cc3e28221de.incomplete'. Error: [Errno 13] Permission denied: '/mnt/tmp_9d7c7c78-c0ca-41b2-a3ef-4c17db21dbd7'. [e2e-llm-inference-service] Continuing without setting permissions. [e2e-llm-inference-service] 2026-07-02 14:11:27.839 1 storage.initializer INFO [kserve_storage.py:download():234] Successfully copied hf://facebook/opt-125m to /mnt/models [e2e-llm-inference-service] 2026-07-02 14:11:27.839 1 storage.initializer INFO [kserve_storage.py:download():235] Model downloaded in 0.4206230500003585 seconds. [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 {"level":"info","ts":1783001488.5412638,"logger":"setup","caller":"runner/runner.go:196","msg":"GIE build","commit-sha":"181aa8358916e19b8844ccc752b2d6153d4b2ad6","build-ref":"v0.9.0-rc.2"} [e2e-llm-inference-service] Flag --model-server-metrics-scheme has been deprecated, This flag is deprecated. Configure via EndpointPickerConfig data layer plugin parameters instead. [e2e-llm-inference-service] {"level":"info","ts":1783001488.541405,"logger":"setup","caller":"runner/runner.go:217","msg":"Flags processed","flags":{"cert-path":"/var/run/kserve/tls","config-file":"","config-text":"apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type: prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n- name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n","disable-endpoint-subset-filter":false,"enable-cert-reload":true,"enable-grpc-stream-metrics":false,"enable-pprof":true,"endpoint-selector":"","endpoint-target-ports":{},"grpc-health-port":9003,"grpc-max-recv-msg-size":"","grpc-max-send-msg-size":"","grpc-port":9002,"ha-enable-leader-election":false,"health-checking":false,"metrics-endpoint-auth":true,"metrics-port":9090,"metrics-staleness-threshold":2000000000,"model-server-metrics-https-insecure-skip-verify":true,"model-server-metrics-path":"/metrics","model-server-metrics-port":0,"model-server-metrics-scheme":"https","pool-group":"inference.networking.k8s.io","pool-name":"tls-verification-test-inference-pool","pool-namespace":"kserve-ci-e2e-test","refresh-metrics-interval":50000000,"refresh-prometheus-metrics-interval":5000000000,"secure-serving":true,"tracing":true,"v":2,"zap-devel":{},"zap-encoder":{},"zap-log-level":{},"zap-stacktrace-level":{},"zap-time-encoding":{}}} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5415034,"logger":"setup.trace","caller":"tracing/telemetry.go:123","msg":"init OTel trace exporter","type":"console"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5418918,"caller":"loader/configloader.go:89","msg":"DEPRECATION: apiVersion inference.networking.x-k8s.io/v1alpha1/EndpointPickerConfig is deprecated","replacement":"llm-d.ai/v1alpha1/EndpointPickerConfig"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5419242,"caller":"loader/configloader.go:121","msg":"Loaded raw configuration","config":"{Plugins: [{Type: single-profile-handler} {Type: queue-scorer} {Type: prefix-cache-scorer} {Type: max-score-picker}], SchedulingProfiles: [{Name: default, Plugins: [{PluginRef: queue-scorer, Weight: 2.00} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: max-score-picker}]}]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5419345,"logger":"setup","caller":"runner/runner.go:622","msg":"Data layer: ENABLED"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.542206,"logger":"setup","caller":"runner/runner.go:281","msg":"Raw config after phase one","config":{"apiVersion":"inference.networking.x-k8s.io/v1alpha1","dataLayer":null,"kind":"EndpointPickerConfig","plugins":[{"name":"single-profile-handler","parameters":null,"type":"single-profile-handler"},{"name":"queue-scorer","parameters":null,"type":"queue-scorer"},{"name":"prefix-cache-scorer","parameters":null,"type":"prefix-cache-scorer"},{"name":"max-score-picker","parameters":null,"type":"max-score-picker"}],"schedulingProfiles":[{"name":"default","plugins":[{"pluginRef":"queue-scorer","weight":2},{"pluginRef":"prefix-cache-scorer","weight":3},{"pluginRef":"max-score-picker","weight":null}]}]}} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5622175,"logger":"utilization-detector/utilization-detector","caller":"utilization/detector.go:83","msg":"Creating new UtilizationDetector","queueDepthThreshold":5,"kvCacheUtilThreshold":0.8,"metricsStalenessThreshold":"200ms","headroom":0} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5623038,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"vllm","mapping":"Mapping{all specs enabled}"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5623376,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"sglang","mapping":"Mapping{disabled: [lora]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.562374,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"trtllm-serve","mapping":"Mapping{disabled: [lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5624323,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"triton-tensorrt-llm","mapping":"Mapping{disabled: [lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5624516,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"triton","mapping":"Mapping{disabled: [kv, lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5624917,"caller":"loader/configloader.go:154","msg":"Instantiated all plugins and applied system defaults. Effective raw configuration","config":"{Plugins: [{Name: single-profile-handler, Type: single-profile-handler} {Name: queue-scorer, Type: queue-scorer} {Name: prefix-cache-scorer, Type: prefix-cache-scorer} {Name: max-score-picker, Type: max-score-picker} {Name: fcfs-ordering-policy, Type: fcfs-ordering-policy} {Name: global-strict-fairness-policy, Type: global-strict-fairness-policy} {Name: static-usage-limit-policy, Type: static-usage-limit-policy} {Name: openai-parser, Type: openai-parser} {Name: anthropic-parser, Type: anthropic-parser} {Name: vllmhttp-parser, Type: vllmhttp-parser} {Name: utilization-detector, Type: utilization-detector} {Name: metrics-data-source, Type: metrics-data-source} {Name: core-metrics-extractor, Type: core-metrics-extractor}], SchedulingProfiles: [{Name: default, Plugins: [{PluginRef: queue-scorer, Weight: 2.00} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: max-score-picker}]}], DataLayer: {Sources: [{PluginRef: metrics-data-source, Extractors: [{PluginRef: core-metrics-extractor}]}], Discovery: }, FlowControl: {MaxBytes: unlimited, MaxRequests: unlimited, SaturationDetector: {PluginRef: utilization-detector}}, RequestHandler: {Parsers: [{PluginRef: openai-parser}, {PluginRef: anthropic-parser}, {PluginRef: vllmhttp-parser}]}}"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5625381,"caller":"approximateprefix/plugin.go:88","msg":"Prefix DataProducer initialized","config":{"autoTune":true,"blockSizeTokens":16,"blockSize":0,"maxPrefixBlocksToMatch":2048,"maxPrefixTokensToMatch":131072,"lruCapacityPerServer":31250}} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5626018,"caller":"approximateprefix/plugin.go:111","msg":"WARNING: configured blockSizeTokens is below the recommended minimum, overriding it.","blockSizeTokens":16,"minimum":64,"issue":"https://github.com/llm-d/llm-d-router/issues/1158"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.562624,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"approx-prefix-cache-producer/approx-prefix-cache-producer","dataKey":"PrefixCacheMatchInfoDataKey/approx-prefix-cache-producer","consumer":"prefix-cache-scorer"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.562661,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"token-producer/token-producer","dataKey":"TokenizedPrompt/token-producer","consumer":"approx-prefix-cache-producer"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5627372,"caller":"runner/runner.go:685","msg":"loaded configuration from file/text successfully"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5627475,"logger":"setup","caller":"runner/runner.go:308","msg":"EPP config after phase two","config":"{SchedulerConfig:{ProfileHandler: single-profile-handler/single-profile-handler, Profiles: map[default:{Filters: [], Scorers: [queue-scorer/queue-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000], Picker: max-score-picker/max-score-picker}]} SaturationDetector:0xc000938580 DataConfig:{Sources:[{Plugin:0xc00097ca20 Extractors:[0xc000938780]}]} FlowControlConfig: ParserRegistry:0xc000938c00}"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5823944,"logger":"setup","caller":"runner/runner.go:352","msg":"Setting pprof handlers"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5824242,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/block"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.582438,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5824428,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/cmdline"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.582447,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/heap"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.582451,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/allocs"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5824554,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/threadcreate"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5824594,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/mutex"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.582464,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/profile"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5824718,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/symbol"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.582477,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/trace"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5824811,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/goroutine"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.582494,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/plugins/state"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5824993,"logger":"setup","caller":"runner/runner.go:373","msg":"parsed config","scheduler-config":"{ProfileHandler: single-profile-handler/single-profile-handler, Profiles: map[default:{Filters: [], Scorers: [queue-scorer/queue-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000], Picker: max-score-picker/max-score-picker}]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5825214,"logger":"setup","caller":"datalayer/runtime.go:99","msg":"Configuring datalayer runtime","numSources":1} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5825286,"logger":"setup","caller":"datalayer/runtime.go:118","msg":"Processing source","source":"metrics-data-source","numExtractors":1} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5825412,"logger":"setup","caller":"datalayer/runtime.go:147","msg":"Source configured","source":"metrics-data-source","extractors":["core-metrics-extractor/core-metrics-extractor"]} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5825596,"logger":"setup","caller":"datalayer/runtime.go:206","msg":"Datalayer runtime configured","pollers":1,"notifiers":0,"endpointSources":0} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5825706,"logger":"setup","caller":"runner/runner.go:833","msg":"Experimental Flow Control layer is disabled, using legacy admission control"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5826457,"logger":"setup","caller":"runner/runner.go:721","msg":"ExtProc server runner added to manager."} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5826561,"logger":"setup","caller":"runner/runner.go:260","msg":"Controller manager starting"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5826845,"logger":"controller-runtime.metrics","caller":"server/server.go:208","msg":"Starting metrics server"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5834394,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"health"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5838883,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"health","port":9003} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5839677,"logger":"controller-runtime.metrics","caller":"server/server.go:247","msg":"Serving metrics server","bindAddress":":9090","secure":false} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5840073,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","source":"kind source: *v1.InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5844812,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"pod","controllerGroup":"","controllerKind":"Pod","source":"kind source: *v1.Pod"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5844576,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite","source":"kind source: *v1alpha2.InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.584993,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"ext-proc"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.585066,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"ext-proc","port":9002} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5852869,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective","source":"kind source: *v1alpha2.InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5899456,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceObjective","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5899472,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceModelRewrite","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.5904038,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.InferencePool","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.592943,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.Pod","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.6859515,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"pod","controllerGroup":"","controllerKind":"Pod"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.6859753,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"pod","controllerGroup":"","controllerKind":"Pod","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783001488.6871073,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.6871254,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783001488.6871436,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.6871614,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783001488.7853806,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1783001488.7854025,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783001488.7855499,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"tls-verification-test-inference-pool","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"tls-verification-test-inference-pool","reconcileID":"d76dde88-6b05-4d9f-93fd-15e52607433c"} [e2e-llm-inference-service] {"level":"info","ts":1783001512.096458,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"tls-verification-test-inference-pool","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"tls-verification-test-inference-pool","reconcileID":"c57e2703-f08d-4ebe-a23f-6cda6cc16b4a"} [e2e-llm-inference-service] {"level":"info","ts":1783001606.6092114,"caller":"controller/pod_reconciler.go:99","msg":"Pod already exists","controller":"pod","controllerGroup":"","controllerKind":"Pod","Pod":{"name":"tls-verification-test-kserve-6d6946dcd6-bdddw","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"tls-verification-test-kserve-6d6946dcd6-bdddw","reconcileID":"6fbfe648-581f-4882-a3c0-b2c8aa9005c4"} [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-service [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 6e2ea6b1-f773-48bd-81cc-a2af8285010b [e2e-llm-inference-service] resourceVersion: '83145' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.132.0.65 [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj [e2e-llm-inference-service] uid: a5e7f68b-3c89-43e5-909d-2c8104bb0150 [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: f4dd1c00-b694-4611-bf2c-ac53557e567c [e2e-llm-inference-service] resourceVersion: '83810' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.132.0.64 [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: tls-verification-test-kserve-6d6946dcd6-bdddw [e2e-llm-inference-service] uid: dc7326bd-f917-466c-922f-739706ba0fda [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-6d6946dcd6-bdddw [e2e-llm-inference-service] generateName: tls-verification-test-kserve-6d6946dcd6- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: dc7326bd-f917-466c-922f-739706ba0fda [e2e-llm-inference-service] resourceVersion: '83807' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 6d6946dcd6 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"10.132.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.132.0.64\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:84:00:40\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: tls-verification-test-kserve-6d6946dcd6 [e2e-llm-inference-service] uid: 617afe42-e508-4967-bf73-581781225c99 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-134-1 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:initContainerStatuses: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.132.0.64"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kube-api-access-2pvc4 [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-2pvc4 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to infer\ [e2e-llm-inference-service] \ RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/* 2>/dev/null\n\ [e2e-llm-inference-service] \ grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/* 2>/dev/null\n\ [e2e-llm-inference-service] \n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"$hca_dir\"\ [e2e-llm-inference-service] \ ]; then\n hca_name=$(basename \"$hca_dir\")\n port_state_file=\"\ [e2e-llm-inference-service] $hca_dir/ports/1/state\" # Assume port 1\n type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\ [e2e-llm-inference-service] \n\n echo \"[Infer RoCE] Check if the port state file ${port_state_file}\ [e2e-llm-inference-service] \ exists and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] &&\ [e2e-llm-inference-service] \ grep -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found active\ [e2e-llm-inference-service] \ HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n else\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Skipping inactive or down HCA: $hca_name\"\ [e2e-llm-inference-service] \n fi\n fi\n done\n\n # Check if we found any active HCAs\n\ [e2e-llm-inference-service] \ if [ ${#active_hcas[@]} -gt 0 ]; then\n # Join the array elements with\ [e2e-llm-inference-service] \ a comma\n hca_port_pairs=()\n for hca in \"${active_hcas[@]}\";\ [e2e-llm-inference-service] \ do\n hca_port_pairs+=(\"${hca}:1\")\n done\n\n active_hca_list=$(IFS=,;\ [e2e-llm-inference-service] \ echo \"${active_hcas[*]}\")\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\"\ [e2e-llm-inference-service] )\n echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n\ [e2e-llm-inference-service] \ export UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found. NCCL_IB_HCA\ [e2e-llm-inference-service] \ will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Finding GID_INDEX for each active HCA (SR-IOV compatible)...\"\ [e2e-llm-inference-service] \n\n # For SR-IOV environments, find the most common IPv4 RoCE v2 GID index\ [e2e-llm-inference-service] \ across all HCAs\n declare -A gid_index_count\n declare -A hca_gid_index\n\ [e2e-llm-inference-service] \n for hca_name in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ Processing HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for\ [e2e-llm-inference-service] \ this HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"$tpath\"\ [e2e-llm-inference-service] \ 2>/dev/null; then\n idx=$(basename \"$tpath\")\n \ [e2e-llm-inference-service] \ gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n \ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo \"\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Found IPv4 RoCE v2 GID for ${hca_name}:\ [e2e-llm-inference-service] \ index=${idx}, gid=${gid_value}\"\n hca_gid_index[\"${hca_name}\"\ [e2e-llm-inference-service] ]=\"${idx}\"\n gid_index_count[\"${idx}\"]=$((${gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]} + 1))\n break # Use first found IPv4 GID per\ [e2e-llm-inference-service] \ HCA\n fi\n fi\n done\n done\n\n\ [e2e-llm-inference-service] \ # Find the most common GID index (most likely to be consistent across\ [e2e-llm-inference-service] \ nodes)\n best_gid_index=\"\"\n max_count=0\n for idx in \"\ [e2e-llm-inference-service] ${!gid_index_count[@]}\"; do\n count=${gid_index_count[\"${idx}\"]}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n \ [e2e-llm-inference-service] \ if [ $count -gt $max_count ]; then\n max_count=$count\n\ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n #\ [e2e-llm-inference-service] \ Use deterministic fallback if tied - prefer index 3 (SR-IOV standard)\n \ [e2e-llm-inference-service] \ if [ ${#gid_index_count[@]} -gt 1 ]; then\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ Multiple GID indices found, selecting most common: ${best_gid_index}\"\n \ [e2e-llm-inference-service] \ # If there's a tie, prefer index 3 as it's most common in SR-IOV setups\n\ [e2e-llm-inference-service] \ if [ -n \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\"\ [e2e-llm-inference-service] \ -eq \"$max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for NCCL,\ [e2e-llm-inference-service] \ NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR: No valid\ [e2e-llm-inference-service] \ IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any HCA.\"\n \ [e2e-llm-inference-service] \ fi\n else\n echo \"[Infer RoCE] No active HCAs found, skipping GID_INDEX\ [e2e-llm-inference-service] \ inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints landed in vLLM\ [e2e-llm-inference-service] \ 0.16.0 (vllm-project/vllm#30011).\n# Older versions still need the blanket\ [e2e-llm-inference-service] \ --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+ ]] &&\ [e2e-llm-inference-service] \ [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort -V | head\ [e2e-llm-inference-service] \ -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout 40\"\ [e2e-llm-inference-service] \nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name \"facebook/opt-125m\"\ [e2e-llm-inference-service] \ \"publishers/kserve-ci-e2e-test/models/facebook/opt-125m\" \\\n --port 8000\ [e2e-llm-inference-service] \ \\\n ${ACCESS_LOG_ARGS} \\\n ${SHUTDOWN_TIMEOUT_ARGS} \\\n --enable-ssl-refresh\ [e2e-llm-inference-service] \ \\\n --ssl-certfile /var/run/kserve/tls/tls.crt \\\n --ssl-keyfile /var/run/kserve/tls/tls.key\ [e2e-llm-inference-service] \ \\\n ${VLLM_ADDITIONAL_ARGS} \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-2pvc4 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: default [e2e-llm-inference-service] serviceAccount: default [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] fsGroup: 1000690000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:11:31Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] hostIP: 10.0.134.1 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.134.1 [e2e-llm-inference-service] podIP: 10.132.0.64 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.132.0.64 [e2e-llm-inference-service] startTime: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] initContainerStatuses: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] state: [e2e-llm-inference-service] terminated: [e2e-llm-inference-service] exitCode: 0 [e2e-llm-inference-service] reason: Completed [e2e-llm-inference-service] startedAt: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] finishedAt: '2026-07-02T14:11:30Z' [e2e-llm-inference-service] containerID: cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8 [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] imageID: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] containerID: cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8 [e2e-llm-inference-service] started: false [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-2pvc4 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T14:11:31Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] imageID: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0 [e2e-llm-inference-service] containerID: cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-2pvc4 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj [e2e-llm-inference-service] generateName: tls-verification-test-kserve-router-scheduler-7f955879f5- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: a5e7f68b-3c89-43e5-909d-2c8104bb0150 [e2e-llm-inference-service] resourceVersion: '83144' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 7f955879f5 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.132.0.65/23"],"mac_address":"0a:58:0a:84:00:41","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.65/23","gateway_ip":"10.132.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.132.0.65\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:84:00:41\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-7f955879f5 [e2e-llm-inference-service] uid: b09b6b35-78d8-4429-9188-7f7739e3ad08 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-134-1 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"b09b6b35-78d8-4429-9188-7f7739e3ad08"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"STORAGE_ALLOW_PATTERNS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:initContainerStatuses: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.132.0.65"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kube-api-access-nvcq4 [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] - name: STORAGE_ALLOW_PATTERNS [e2e-llm-inference-service] value: '["tokenizer.json", "tokenizer_config.json", "special_tokens_map.json", [e2e-llm-inference-service] "vocab.json", "merges.txt", "config.json", "generation_config.json"]' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-nvcq4 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - tls-verification-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type: prefix-cache-scorer\n\ [e2e-llm-inference-service] - type: max-score-picker\nschedulingProfiles:\n- name: default\n plugins:\n\ [e2e-llm-inference-service] \ - pluginRef: queue-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n\ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-nvcq4 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: tls-verification-test-epp-sa [e2e-llm-inference-service] serviceAccount: tls-verification-test-epp-sa [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] fsGroup: 1000690000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: tls-verification-test-epp-sa-dockercfg-lg5ls [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:11:28Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] hostIP: 10.0.134.1 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.134.1 [e2e-llm-inference-service] podIP: 10.132.0.65 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.132.0.65 [e2e-llm-inference-service] startTime: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] initContainerStatuses: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] state: [e2e-llm-inference-service] terminated: [e2e-llm-inference-service] exitCode: 0 [e2e-llm-inference-service] reason: Completed [e2e-llm-inference-service] startedAt: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] finishedAt: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] containerID: cri-o://d9461b9905771918ca854c7298a769f287a9c99bed534aa1a989063d165d9e38 [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] imageID: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] containerID: cri-o://d9461b9905771918ca854c7298a769f287a9c99bed534aa1a989063d165d9e38 [e2e-llm-inference-service] started: false [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] - name: kube-api-access-nvcq4 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T14:11:28Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] imageID: ghcr.io/llm-d/llm-d-router-endpoint-picker@sha256:06b6c75d77afd0e07053402752a9736c2dfbc12a306d0d37d963aac4c1d4e6a6 [e2e-llm-inference-service] containerID: cri-o://06a85f783d00b0ab049cfe211794a4e6e2c9d99ab5b367e8a81dec0d1e949da6 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-nvcq4 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-sa [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: eb3074a7-8028-4ce2-8eb1-8d805ebeb1b3 [e2e-llm-inference-service] resourceVersion: '82699' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: tls-verification-test-epp-sa-dockercfg-lg5ls [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"tls-verification-test-epp-sa-dockercfg-lg5ls"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-x9l7t"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: tls-verification-test-epp-sa-dockercfg-lg5ls [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: tls-verification-test-epp-sa-dockercfg-lg5ls [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-service [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 967dce25-64b6-4ec3-b0a1-15ce6c104f6f [e2e-llm-inference-service] resourceVersion: '82724' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] targetPort: grpc [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] targetPort: grpc-health [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] targetPort: metrics [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] targetPort: zmq [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] clusterIP: 172.31.159.44 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.159.44 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 881cf2c2-891b-47bc-8df3-c774062e46f1 [e2e-llm-inference-service] resourceVersion: '82691' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] targetPort: 8000 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] clusterIP: 172.31.8.21 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.8.21 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 09352d53-4084-4beb-9bcb-2d224e0d3fbe [e2e-llm-inference-service] resourceVersion: '83813' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:rollingUpdate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:maxSurge: {} [e2e-llm-inference-service] f:maxUnavailable: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to\ [e2e-llm-inference-service] \ infer RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/*\ [e2e-llm-inference-service] \ 2>/dev/null\n grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/*\ [e2e-llm-inference-service] \ 2>/dev/null\n\n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"\ [e2e-llm-inference-service] $hca_dir\" ]; then\n hca_name=$(basename \"$hca_dir\")\n \ [e2e-llm-inference-service] \ port_state_file=\"$hca_dir/ports/1/state\" # Assume port 1\n \ [e2e-llm-inference-service] \ type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] Check if the port state file ${port_state_file} exists\ [e2e-llm-inference-service] \ and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] && grep\ [e2e-llm-inference-service] \ -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found\ [e2e-llm-inference-service] \ active HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n\ [e2e-llm-inference-service] \ else\n echo \"[Infer RoCE] Skipping inactive or\ [e2e-llm-inference-service] \ down HCA: $hca_name\"\n fi\n fi\n done\n\n # Check if\ [e2e-llm-inference-service] \ we found any active HCAs\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ # Join the array elements with a comma\n hca_port_pairs=()\n \ [e2e-llm-inference-service] \ for hca in \"${active_hcas[@]}\"; do\n hca_port_pairs+=(\"\ [e2e-llm-inference-service] ${hca}:1\")\n done\n\n active_hca_list=$(IFS=,; echo \"${active_hcas[*]}\"\ [e2e-llm-inference-service] )\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n export\ [e2e-llm-inference-service] \ UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found.\ [e2e-llm-inference-service] \ NCCL_IB_HCA will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt\ [e2e-llm-inference-service] \ 0 ]; then\n echo \"[Infer RoCE] Finding GID_INDEX for each active\ [e2e-llm-inference-service] \ HCA (SR-IOV compatible)...\"\n\n # For SR-IOV environments, find\ [e2e-llm-inference-service] \ the most common IPv4 RoCE v2 GID index across all HCAs\n declare\ [e2e-llm-inference-service] \ -A gid_index_count\n declare -A hca_gid_index\n\n for hca_name\ [e2e-llm-inference-service] \ in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE] Processing\ [e2e-llm-inference-service] \ HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for this\ [e2e-llm-inference-service] \ HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"\ [e2e-llm-inference-service] $tpath\" 2>/dev/null; then\n idx=$(basename \"$tpath\"\ [e2e-llm-inference-service] )\n gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n\ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo\ [e2e-llm-inference-service] \ \"\")\n echo \"[Infer RoCE] Found IPv4 RoCE v2 GID\ [e2e-llm-inference-service] \ for ${hca_name}: index=${idx}, gid=${gid_value}\"\n \ [e2e-llm-inference-service] \ hca_gid_index[\"${hca_name}\"]=\"${idx}\"\n gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]=$((${gid_index_count[\"${idx}\"]} + 1))\n \ [e2e-llm-inference-service] \ break # Use first found IPv4 GID per HCA\n fi\n \ [e2e-llm-inference-service] \ fi\n done\n done\n\n # Find the most common\ [e2e-llm-inference-service] \ GID index (most likely to be consistent across nodes)\n best_gid_index=\"\ [e2e-llm-inference-service] \"\n max_count=0\n for idx in \"${!gid_index_count[@]}\"; do\n\ [e2e-llm-inference-service] \ count=${gid_index_count[\"${idx}\"]}\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n if [ $count\ [e2e-llm-inference-service] \ -gt $max_count ]; then\n max_count=$count\n \ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n # Use deterministic\ [e2e-llm-inference-service] \ fallback if tied - prefer index 3 (SR-IOV standard)\n if [ ${#gid_index_count[@]}\ [e2e-llm-inference-service] \ -gt 1 ]; then\n echo \"[Infer RoCE] Multiple GID indices found,\ [e2e-llm-inference-service] \ selecting most common: ${best_gid_index}\"\n # If there's a tie,\ [e2e-llm-inference-service] \ prefer index 3 as it's most common in SR-IOV setups\n if [ -n\ [e2e-llm-inference-service] \ \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\" -eq \"\ [e2e-llm-inference-service] $max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\"\ [e2e-llm-inference-service] \ ]; then\n echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \n echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for\ [e2e-llm-inference-service] \ NCCL, NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR:\ [e2e-llm-inference-service] \ No valid IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any\ [e2e-llm-inference-service] \ HCA.\"\n fi\n else\n echo \"[Infer RoCE] No active HCAs found,\ [e2e-llm-inference-service] \ skipping GID_INDEX inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ landed in vLLM 0.16.0 (vllm-project/vllm#30011).\n# Older versions still\ [e2e-llm-inference-service] \ need the blanket --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+\ [e2e-llm-inference-service] \ ]] && [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort\ [e2e-llm-inference-service] \ -V | head -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout\ [e2e-llm-inference-service] \ 40\"\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name\ [e2e-llm-inference-service] \ \"facebook/opt-125m\" \"publishers/kserve-ci-e2e-test/models/facebook/opt-125m\"\ [e2e-llm-inference-service] \ \\\n --port 8000 \\\n ${ACCESS_LOG_ARGS} \\\n ${SHUTDOWN_TIMEOUT_ARGS}\ [e2e-llm-inference-service] \ \\\n --enable-ssl-refresh \\\n --ssl-certfile /var/run/kserve/tls/tls.crt\ [e2e-llm-inference-service] \ \\\n --ssl-keyfile /var/run/kserve/tls/tls.key \\\n ${VLLM_ADDITIONAL_ARGS}\ [e2e-llm-inference-service] \ \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: RollingUpdate [e2e-llm-inference-service] rollingUpdate: [e2e-llm-inference-service] maxUnavailable: 25% [e2e-llm-inference-service] maxSurge: 25% [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "tls-verification-test-kserve-6d6946dcd6" has successfully [e2e-llm-inference-service] progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: aa08932a-9d9e-4557-8067-1fc15ff757d9 [e2e-llm-inference-service] resourceVersion: '83148' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"STORAGE_ALLOW_PATTERNS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] - name: STORAGE_ALLOW_PATTERNS [e2e-llm-inference-service] value: '["tokenizer.json", "tokenizer_config.json", "special_tokens_map.json", [e2e-llm-inference-service] "vocab.json", "merges.txt", "config.json", "generation_config.json"]' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - tls-verification-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n-\ [e2e-llm-inference-service] \ name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: tls-verification-test-epp-sa [e2e-llm-inference-service] serviceAccount: tls-verification-test-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: Recreate [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "tls-verification-test-kserve-router-scheduler-7f955879f5" [e2e-llm-inference-service] has successfully progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-6d6946dcd6 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 617afe42-e508-4967-bf73-581781225c99 [e2e-llm-inference-service] resourceVersion: '83812' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 6d6946dcd6 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '2' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: tls-verification-test-kserve [e2e-llm-inference-service] uid: 09352d53-4084-4beb-9bcb-2d224e0d3fbe [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"09352d53-4084-4beb-9bcb-2d224e0d3fbe"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"TORCHINDUCTOR_CACHE_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"USER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_CPU_KVCACHE_SPACE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 6d6946dcd6 [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 6d6946dcd6 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/bash [e2e-llm-inference-service] - -c [e2e-llm-inference-service] - "if [ -f /etc/profile.d/ibm-aiu-setup.sh ]; then\n source /etc/profile.d/ibm-aiu-setup.sh\n\ [e2e-llm-inference-service] fi\n\nif [ \"$KSERVE_INFER_ROCE\" = \"true\" ]; then\n echo \"Trying to\ [e2e-llm-inference-service] \ infer RoCE configs ... \"\n grep -H . /sys/class/infiniband/*/ports/*/gids/*\ [e2e-llm-inference-service] \ 2>/dev/null\n grep -H . /sys/class/infiniband/*/ports/*/gid_attrs/types/*\ [e2e-llm-inference-service] \ 2>/dev/null\n\n cat /proc/driver/nvidia/params\n\n KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-\"\ [e2e-llm-inference-service] RoCE v2\"}\n\n echo \"[Infer RoCE] Discovering active HCAs ...\"\n active_hcas=()\n\ [e2e-llm-inference-service] \ # Loop through all mlx5 devices found in sysfs\n for hca_dir in /sys/class/infiniband/mlx5_*;\ [e2e-llm-inference-service] \ do\n # Ensure it's a directory before proceeding\n if [ -d \"\ [e2e-llm-inference-service] $hca_dir\" ]; then\n hca_name=$(basename \"$hca_dir\")\n \ [e2e-llm-inference-service] \ port_state_file=\"$hca_dir/ports/1/state\" # Assume port 1\n \ [e2e-llm-inference-service] \ type_file=\"$hca_dir/ports/1/gid_attrs/types/*\"\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] Check if the port state file ${port_state_file} exists\ [e2e-llm-inference-service] \ and contains 'ACTIVE'\"\n if [ -f \"$port_state_file\" ] && grep\ [e2e-llm-inference-service] \ -q \"ACTIVE\" \"$port_state_file\" && grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\"\ [e2e-llm-inference-service] \ ${type_file} 2>/dev/null; then\n echo \"[Infer RoCE] Found\ [e2e-llm-inference-service] \ active HCA: $hca_name\"\n active_hcas+=(\"$hca_name\")\n\ [e2e-llm-inference-service] \ else\n echo \"[Infer RoCE] Skipping inactive or\ [e2e-llm-inference-service] \ down HCA: $hca_name\"\n fi\n fi\n done\n\n # Check if\ [e2e-llm-inference-service] \ we found any active HCAs\n if [ ${#active_hcas[@]} -gt 0 ]; then\n \ [e2e-llm-inference-service] \ # Join the array elements with a comma\n hca_port_pairs=()\n \ [e2e-llm-inference-service] \ for hca in \"${active_hcas[@]}\"; do\n hca_port_pairs+=(\"\ [e2e-llm-inference-service] ${hca}:1\")\n done\n\n active_hca_list=$(IFS=,; echo \"${active_hcas[*]}\"\ [e2e-llm-inference-service] )\n hca_port_pairs_list=$(IFS=,; echo \"${hca_port_pairs[*]}\")\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Setting active HCAs: ${active_hca_list}\"\n \ [e2e-llm-inference-service] \ export NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n export\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n export\ [e2e-llm-inference-service] \ UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n\n echo\ [e2e-llm-inference-service] \ \"[Infer RoCE] NCCL_IB_HCA=${NCCL_IB_HCA}\"\n echo \"[Infer RoCE]\ [e2e-llm-inference-service] \ NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}\"\n echo \"[Infer RoCE] UCX_NET_DEVICES=${UCX_NET_DEVICES}\"\ [e2e-llm-inference-service] \n else\n echo \"[Infer RoCE] WARNING: No active RoCE HCAs found.\ [e2e-llm-inference-service] \ NCCL_IB_HCA will not be set.\"\n fi\n\n if [ ${#active_hcas[@]} -gt\ [e2e-llm-inference-service] \ 0 ]; then\n echo \"[Infer RoCE] Finding GID_INDEX for each active\ [e2e-llm-inference-service] \ HCA (SR-IOV compatible)...\"\n\n # For SR-IOV environments, find\ [e2e-llm-inference-service] \ the most common IPv4 RoCE v2 GID index across all HCAs\n declare\ [e2e-llm-inference-service] \ -A gid_index_count\n declare -A hca_gid_index\n\n for hca_name\ [e2e-llm-inference-service] \ in \"${active_hcas[@]}\"; do\n echo \"[Infer RoCE] Processing\ [e2e-llm-inference-service] \ HCA: ${hca_name}\"\n\n # Find all RoCE v2 IPv4 GIDs for this\ [e2e-llm-inference-service] \ HCA and count by index\n for tpath in /sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*;\ [e2e-llm-inference-service] \ do\n if grep -q \"${KSERVE_INFER_IB_GID_INDEX_GREP}\" \"\ [e2e-llm-inference-service] $tpath\" 2>/dev/null; then\n idx=$(basename \"$tpath\"\ [e2e-llm-inference-service] )\n gid_file=\"/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}\"\ [e2e-llm-inference-service] \n # Check for IPv4 GID (contains ffff:)\n \ [e2e-llm-inference-service] \ if [ -f \"$gid_file\" ] && grep -q \"ffff:\" \"$gid_file\"; then\n\ [e2e-llm-inference-service] \ gid_value=$(cat \"$gid_file\" 2>/dev/null || echo\ [e2e-llm-inference-service] \ \"\")\n echo \"[Infer RoCE] Found IPv4 RoCE v2 GID\ [e2e-llm-inference-service] \ for ${hca_name}: index=${idx}, gid=${gid_value}\"\n \ [e2e-llm-inference-service] \ hca_gid_index[\"${hca_name}\"]=\"${idx}\"\n gid_index_count[\"\ [e2e-llm-inference-service] ${idx}\"]=$((${gid_index_count[\"${idx}\"]} + 1))\n \ [e2e-llm-inference-service] \ break # Use first found IPv4 GID per HCA\n fi\n \ [e2e-llm-inference-service] \ fi\n done\n done\n\n # Find the most common\ [e2e-llm-inference-service] \ GID index (most likely to be consistent across nodes)\n best_gid_index=\"\ [e2e-llm-inference-service] \"\n max_count=0\n for idx in \"${!gid_index_count[@]}\"; do\n\ [e2e-llm-inference-service] \ count=${gid_index_count[\"${idx}\"]}\n echo \"[Infer\ [e2e-llm-inference-service] \ RoCE] GID_INDEX ${idx} found on ${count} HCAs\"\n if [ $count\ [e2e-llm-inference-service] \ -gt $max_count ]; then\n max_count=$count\n \ [e2e-llm-inference-service] \ best_gid_index=\"$idx\"\n fi\n done\n\n # Use deterministic\ [e2e-llm-inference-service] \ fallback if tied - prefer index 3 (SR-IOV standard)\n if [ ${#gid_index_count[@]}\ [e2e-llm-inference-service] \ -gt 1 ]; then\n echo \"[Infer RoCE] Multiple GID indices found,\ [e2e-llm-inference-service] \ selecting most common: ${best_gid_index}\"\n # If there's a tie,\ [e2e-llm-inference-service] \ prefer index 3 as it's most common in SR-IOV setups\n if [ -n\ [e2e-llm-inference-service] \ \"${gid_index_count['3']}\" ] && [ \"${gid_index_count['3']}\" -eq \"\ [e2e-llm-inference-service] $max_count\" ]; then\n best_gid_index=\"3\"\n \ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using deterministic fallback: GID_INDEX=3 (SR-IOV\ [e2e-llm-inference-service] \ standard)\"\n fi\n fi\n\n # Check if GID_INDEX is already\ [e2e-llm-inference-service] \ set via environment variables\n if [ -n \"${NCCL_IB_GID_INDEX}\"\ [e2e-llm-inference-service] \ ]; then\n echo \"[Infer RoCE] Using pre-configured NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ from environment\"\n export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Using pre-configured GID_INDEX=${NCCL_IB_GID_INDEX}\ [e2e-llm-inference-service] \ for NCCL, NVSHMEM, and UCX\"\n elif [ -n \"$best_gid_index\" ]; then\n\ [e2e-llm-inference-service] \ echo \"[Infer RoCE] Selected GID_INDEX: ${best_gid_index} (found\ [e2e-llm-inference-service] \ on ${max_count} HCAs)\"\n\n export NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \ export UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n\ [e2e-llm-inference-service] \n echo \"[Infer RoCE] Exported GID_INDEX=${best_gid_index} for\ [e2e-llm-inference-service] \ NCCL, NVSHMEM, and UCX\"\n else\n echo \"[Infer RoCE] ERROR:\ [e2e-llm-inference-service] \ No valid IPv4 ${KSERVE_INFER_IB_GID_INDEX_GREP} GID_INDEX found on any\ [e2e-llm-inference-service] \ HCA.\"\n fi\n else\n echo \"[Infer RoCE] No active HCAs found,\ [e2e-llm-inference-service] \ skipping GID_INDEX inference.\"\n fi\nfi\n\n# --disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ landed in vLLM 0.16.0 (vllm-project/vllm#30011).\n# Older versions still\ [e2e-llm-inference-service] \ need the blanket --disable-uvicorn-access-log.\nACCESS_LOG_ARGS=\"--disable-uvicorn-access-log\"\ [e2e-llm-inference-service] \nVLLM_VERSION=$(vllm --version 2>/dev/null | tail -1 | awk '{print $NF}')\n\ [e2e-llm-inference-service] echo \"[access-log-detect] vllm version='${VLLM_VERSION}'\"\nif [[ \"$VLLM_VERSION\"\ [e2e-llm-inference-service] \ =~ ^[0-9]+\\.[0-9]+ ]] && [ \"$(printf '%s\\n%s\\n' \"0.16.0\" \"${VLLM_VERSION}\"\ [e2e-llm-inference-service] \ | sort -V | head -1)\" = \"0.16.0\" ]; then\n ACCESS_LOG_ARGS=\"--disable-access-log-for-endpoints\ [e2e-llm-inference-service] \ /health,/metrics,/ping\"\nfi\necho \"[access-log-detect] selected ACCESS_LOG_ARGS='${ACCESS_LOG_ARGS}'\"\ [e2e-llm-inference-service] \n\n# --shutdown-timeout landed in vLLM 0.18.0 (vllm-project/vllm#36666).\n\ [e2e-llm-inference-service] SHUTDOWN_TIMEOUT_ARGS=\"\"\nif [[ \"$VLLM_VERSION\" =~ ^[0-9]+\\.[0-9]+\ [e2e-llm-inference-service] \ ]] && [ \"$(printf '%s\\n%s\\n' \"0.18.0\" \"${VLLM_VERSION}\" | sort\ [e2e-llm-inference-service] \ -V | head -1)\" = \"0.18.0\" ]; then\n SHUTDOWN_TIMEOUT_ARGS=\"--shutdown-timeout\ [e2e-llm-inference-service] \ 40\"\nfi\n\neval \"exec vllm serve /mnt/models \\\n --served-model-name\ [e2e-llm-inference-service] \ \"facebook/opt-125m\" \"publishers/kserve-ci-e2e-test/models/facebook/opt-125m\"\ [e2e-llm-inference-service] \ \\\n --port 8000 \\\n ${ACCESS_LOG_ARGS} \\\n ${SHUTDOWN_TIMEOUT_ARGS}\ [e2e-llm-inference-service] \ \\\n --enable-ssl-refresh \\\n --ssl-certfile /var/run/kserve/tls/tls.crt\ [e2e-llm-inference-service] \ \\\n --ssl-keyfile /var/run/kserve/tls/tls.key \\\n ${VLLM_ADDITIONAL_ARGS}\ [e2e-llm-inference-service] \ \\\n $@\"" [e2e-llm-inference-service] - -- [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: DEBUG [e2e-llm-inference-service] - name: VLLM_CPU_KVCACHE_SPACE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: VLLM_ENABLE_V1_MULTIPROCESSING [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: USER [e2e-llm-inference-service] value: nonroot [e2e-llm-inference-service] - name: TORCHINDUCTOR_CACHE_DIR [e2e-llm-inference-service] value: /tmp/torchinductor-cache [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '2' [e2e-llm-inference-service] memory: 7Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-7f955879f5 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: b09b6b35-78d8-4429-9188-7f7739e3ad08 [e2e-llm-inference-service] resourceVersion: '83147' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 7f955879f5 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler [e2e-llm-inference-service] uid: aa08932a-9d9e-4557-8067-1fc15ff757d9 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"aa08932a-9d9e-4557-8067-1fc15ff757d9"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:initContainers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"storage-initializer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"AWS_ACCESS_KEY_ID"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_ENDPOINT_URL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"AWS_SECRET_ACCESS_KEY"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:valueFrom: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:secretKeyRef: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_HIGH_PERFORMANCE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_ENDPOINT"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_USE_HTTPS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"S3_VERIFY_SSL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"STORAGE_ALLOW_PATTERNS"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/mnt/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"kserve-provision-location"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 7f955879f5 [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 7f955879f5 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] initContainers: [e2e-llm-inference-service] - name: storage-initializer [e2e-llm-inference-service] image: quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b [e2e-llm-inference-service] args: [e2e-llm-inference-service] - hf://facebook/opt-125m [e2e-llm-inference-service] - /mnt/models [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_ACCESS_KEY_ID [e2e-llm-inference-service] - name: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] valueFrom: [e2e-llm-inference-service] secretKeyRef: [e2e-llm-inference-service] name: seaweedfs-s3-creds [e2e-llm-inference-service] key: AWS_SECRET_ACCESS_KEY [e2e-llm-inference-service] - name: S3_USE_HTTPS [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: S3_ENDPOINT [e2e-llm-inference-service] value: s3-service.kserve:8333 [e2e-llm-inference-service] - name: AWS_ENDPOINT_URL [e2e-llm-inference-service] value: http://s3-service.kserve:8333 [e2e-llm-inference-service] - name: S3_VERIFY_SSL [e2e-llm-inference-service] value: '0' [e2e-llm-inference-service] - name: AWS_CA_BUNDLE [e2e-llm-inference-service] value: /etc/ssl/custom-certs/cabundle.crt [e2e-llm-inference-service] - name: AWS_CA_BUNDLE_CONFIGMAP [e2e-llm-inference-service] value: odh-kserve-custom-ca-bundle [e2e-llm-inference-service] - name: HF_HUB_ENABLE_HF_TRANSFER [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_HIGH_PERFORMANCE [e2e-llm-inference-service] value: '1' [e2e-llm-inference-service] - name: HF_XET_NUM_CONCURRENT_RANGE_GETS [e2e-llm-inference-service] value: '8' [e2e-llm-inference-service] - name: STORAGE_ALLOW_PATTERNS [e2e-llm-inference-service] value: '["tokenizer.json", "tokenizer_config.json", "special_tokens_map.json", [e2e-llm-inference-service] "vocab.json", "merges.txt", "config.json", "generation_config.json"]' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 24Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 100m [e2e-llm-inference-service] memory: 100Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: kserve-provision-location [e2e-llm-inference-service] mountPath: /mnt/models [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - tls-verification-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n-\ [e2e-llm-inference-service] \ name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: tls-verification-test-epp-sa [e2e-llm-inference-service] serviceAccount: tls-verification-test-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-rb [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 4d106265-7b40-4015-8b68-2bd2786f28c0 [e2e-llm-inference-service] resourceVersion: '82718' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: tls-verification-test-epp-sa [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] name: tls-verification-test-epp-role [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-role [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: d4cc6b84-3a4b-4f67-9699-cee869bdc5f6 [e2e-llm-inference-service] resourceVersion: '82715' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] - create [e2e-llm-inference-service] - update [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - delete [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-service-d4x2h [e2e-llm-inference-service] generateName: tls-verification-test-epp-service- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 188bdafa-21cb-4222-a066-30eed4d3b45b [e2e-llm-inference-service] resourceVersion: '83146' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: tls-verification-test-epp-service [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-epp-service [e2e-llm-inference-service] uid: 967dce25-64b6-4ec3-b0a1-15ce6c104f6f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:12:00Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"967dce25-64b6-4ec3-b0a1-15ce6c104f6f"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.132.0.65 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj [e2e-llm-inference-service] uid: a5e7f68b-3c89-43e5-909d-2c8104bb0150 [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc-xqdkt [e2e-llm-inference-service] generateName: tls-verification-test-kserve-workload-svc- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 559e59ce-4393-4099-8a39-1b6f4362ceea [e2e-llm-inference-service] resourceVersion: '83809' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] uid: 881cf2c2-891b-47bc-8df3-c774062e46f1 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"881cf2c2-891b-47bc-8df3-c774062e46f1"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.132.0.64 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: tls-verification-test-kserve-6d6946dcd6-bdddw [e2e-llm-inference-service] uid: dc7326bd-f917-466c-922f-739706ba0fda [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-rb [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 4d106265-7b40-4015-8b68-2bd2786f28c0 [e2e-llm-inference-service] resourceVersion: '82718' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:kserve-ci-e2e-test:tls-verification-test-epp-sa [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: tls-verification-test-epp-sa [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: tls-verification-test-epp-role [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-epp-role [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: d4cc6b84-3a4b-4f67-9699-cee869bdc5f6 [e2e-llm-inference-service] resourceVersion: '82715' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - create [e2e-llm-inference-service] - delete [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - update [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-route [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '83080' [e2e-llm-inference-service] uid: 3af23dba-c3af-4944-a4d9-70f069493ab2 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/tls-verification-test/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/tls-verification-test/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/tls-verification-test/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/tls-verification-test/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/tls-verification-test [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [kserve-ci-e2e-test/tls-verification-test-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-route [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '83080' [e2e-llm-inference-service] uid: 3af23dba-c3af-4944-a4d9-70f069493ab2 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/tls-verification-test/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/tls-verification-test/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/tls-verification-test/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/tls-verification-test/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/tls-verification-test [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [kserve-ci-e2e-test/tls-verification-test-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:endpointPickerRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:number: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:matchLabels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPorts: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '83069' [e2e-llm-inference-service] uid: 0cfbf6a6-3d93-46a5-b499-57507d9fc351 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] appProtocol: http [e2e-llm-inference-service] endpointPickerRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-epp-service [e2e-llm-inference-service] port: [e2e-llm-inference-service] number: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPorts: [e2e-llm-inference-service] - number: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] message: Referenced by an HTTPRoute accepted by the parentRef Gateway [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] message: Referenced ExtensionRef resolved successfully [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: networking.istio.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] kind: AuthPolicy [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-policies [e2e-llm-inference-service] app.kubernetes.io/managed-by: odh-model-controller [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:rules: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:authentication: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:public: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:anonymous: {} [e2e-llm-inference-service] f:credentials: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:overrides: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fairness: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:response: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:success: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:headers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:x-gateway-inference-fairness-id: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:x-gateway-inference-objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:targetRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Accepted"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Enforced"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:11:29Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-route-authn [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '82868' [e2e-llm-inference-service] uid: d5d433e0-a2f3-43cb-b702-04f4f2a3f428 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] rules: [e2e-llm-inference-service] authentication: [e2e-llm-inference-service] public: [e2e-llm-inference-service] anonymous: {} [e2e-llm-inference-service] credentials: {} [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] overrides: [e2e-llm-inference-service] fairness: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] objective: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] response: [e2e-llm-inference-service] success: [e2e-llm-inference-service] headers: [e2e-llm-inference-service] x-gateway-inference-fairness-id: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.fairness [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] x-gateway-inference-objective: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.objective [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] name: tls-verification-test-kserve-route [e2e-llm-inference-service] status: [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:28Z' [e2e-llm-inference-service] message: AuthPolicy has been accepted [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:11:29Z' [e2e-llm-inference-service] message: AuthPolicy has been successfully enforced [e2e-llm-inference-service] reason: Enforced [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Enforced [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '82756' [e2e-llm-inference-service] uid: ece75f51-61d1-48f7-8051-1bdc4a5da6a7 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '83075' [e2e-llm-inference-service] uid: 172ed0a5-e597-4e51-880c-97ffe7703257 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-inference-pool-ip-e386aca5.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '82797' [e2e-llm-inference-service] uid: 91da174d-3c41-42c9-8b55-08f7493559ab [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '82756' [e2e-llm-inference-service] uid: ece75f51-61d1-48f7-8051-1bdc4a5da6a7 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '83075' [e2e-llm-inference-service] uid: 172ed0a5-e597-4e51-880c-97ffe7703257 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-inference-pool-ip-e386aca5.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '82797' [e2e-llm-inference-service] uid: 91da174d-3c41-42c9-8b55-08f7493559ab [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '82756' [e2e-llm-inference-service] uid: ece75f51-61d1-48f7-8051-1bdc4a5da6a7 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:52Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '83075' [e2e-llm-inference-service] uid: 172ed0a5-e597-4e51-880c-97ffe7703257 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-inference-pool-ip-e386aca5.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:27Z' [e2e-llm-inference-service] name: tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '82797' [e2e-llm-inference-service] uid: 91da174d-3c41-42c9-8b55-08f7493559ab [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: tls-verification-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: tls-verification-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"1b5ef653-070a-4fb2-a295-fd23a54e959b"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:extensionRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:portNumber: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPortNumber: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:11:26Z' [e2e-llm-inference-service] name: tls-verification-test-inference-pool [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: tls-verification-test [e2e-llm-inference-service] uid: 1b5ef653-070a-4fb2-a295-fd23a54e959b [e2e-llm-inference-service] resourceVersion: '82738' [e2e-llm-inference-service] uid: e183881d-9234-408d-b20f-fcb15161c287 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] extensionRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: tls-verification-test-epp-service [e2e-llm-inference-service] portNumber: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPortNumber: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parent: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '1970-01-01T00:00:00Z' [e2e-llm-inference-service] message: Waiting for controller [e2e-llm-inference-service] reason: Pending [e2e-llm-inference-service] status: Unknown [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Status [e2e-llm-inference-service] name: default [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-6d6946dcd6-bdddw [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:28:33Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 6d6946dcd6 [e2e-llm-inference-service] timestamp: '2026-07-02T14:28:20Z' [e2e-llm-inference-service] window: 18.769s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 104601417n [e2e-llm-inference-service] memory: 2403080Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:28:33Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: tls-verification-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 7f955879f5 [e2e-llm-inference-service] timestamp: '2026-07-02T14:28:13Z' [e2e-llm-inference-service] window: 17.817s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 56527922n [e2e-llm-inference-service] memory: 32108Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [delete_llmisvc] [2026-07-02T14:28:34.199374] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'tls-verification-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-tls-verification-86160844'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-tls-verific-00217f09'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-tls-verificat-7c5fca28'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 2 pod(s) for tls-verification-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': [e2e-llm-inference-service] {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n [e2e-llm-inference-service] 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s. [e2e-llm-inference-service] v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n [e2e-llm-inference-service] 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '83807',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n [e2e-llm-inference-service] 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n [e2e-llm-inference-service] '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n [e2e-llm-inference-service] 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n [e2e-llm-inference-service] 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n [e2e-llm-inference-service] ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n [e2e-llm-inference-service] 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n [e2e-llm-inference-service] 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n [e2e-llm-inference-service] 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n [e2e-llm-inference-service] 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n [e2e-llm-inference-service] 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None [e2e-llm-inference-service] ,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n [e2e-llm-inference-service] 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n [e2e-llm-inference-service] 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n [e2e-llm-inference-service] 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e604223 [e2e-llm-inference-service] 65f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}, {'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'app.kubernetes.io/version': '0.9.0',\n 'certificates.kserve.io/expiration-v2': 'true',\n 'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.65/23"],"mac_address":"0a:58:0a:84:00:41","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_add [e2e-llm-inference-service] ress":"10.132.0.65/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.65"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:41",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-router-scheduler-7f955879f5-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-router-scheduler',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'pod-template-hash': '7f955879f5'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'.': {},\n 'f:app.kubernetes.io/version': {},\n 'f:certificates.kserve.io/expiration-v2': {}},\n 'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n [e2e-llm-inference-service] 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"b09b6b35-78d8-4429-9188-7f7739e3ad08"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"SSL_CERT_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:grpc': {'.': {},\n 'f:port': {},\n 'f:service': {}},\n 'f:initialDelaySeconds': {},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":5557,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:protocol': {}},\n 'k:{"containerPort":9002,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n 'f:protocol': {}},\n 'k:{"containerPort":9003,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n 'f:protocol': {}},\n 'k:{"containerPort":9090,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:grpc': {'.': {},\n 'f:port': {},\n 'f:service': {}},\n 'f:initialDelaySeconds': {},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n [e2e-llm-inference-service] 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/tmp/tokenizer"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n [e2e-llm-inference-service] 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}, [e2e-llm-inference-service] \n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"STORAGE_ALLOW_PATTERNS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n [e2e-llm-inference-service] 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:serviceAccount': {},\n 'f:serviceAccountName': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tokenizer-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tokenizer-tmp"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tokenizer-uds"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal())},\n {'api_version': 'v1',\n [e2e-llm-inference-service] 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0. [e2e-llm-inference-service] 65"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-router-scheduler-7f955879f5',\n 'uid': 'b09b6b35-78d8-4429-9188-7f7739e3ad08'}],\n 'resource_version': '83144',\n 'self_link': None,\n 'uid': 'a5e7f68b-3c89-43e5-909d-2c8104bb0150'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--config-text',\n 'apiVersion: '\n 'inference.networking.x-k8s.io/v1alpha1\n'\n 'kind: EndpointPickerConfig\n'\n 'plugins:\n'\n '- type: single-profile-handler\n'\n '- type: queue-scorer\n'\n '- type: prefix-cache-scorer\n'\n '- type: max-score-picker\n'\n 'schedulingProfiles:\n'\n '- name: default\n'\n ' plugins:\n'\n ' - pluginRef: queue-scorer\n'\n ' weight: 2\n'\n ' - pluginRef: prefix-cache-scorer\n'\n ' weight: 3\n'\n ' - pluginRef: max-score-picker\n'],\n 'command': ['/app/epp',\n '--pool-name',\n 'tls-verification-test-inference-pool',\n '--pool-namespace',\n 'kserve-ci-e2e-test',\n '--zap-encoder',\n 'json',\n '--grpc-port',\n '9002',\n '--grpc-health-port',\n '9003',\n '--enable-cert-reload=true',\n '--secure-serving=true',\n '--model-server-metrics-scheme=https',\n '--cert-path=/var/run/kserve/tls'],\n 'env': [{'name': 'SSL_CERT_DIR',\n 'value': '/var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n [e2e-llm-inference-service] 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 3,\n 'grpc': {'port': 9003,\n 'service': 'liveness'},\n 'http_get': None,\n 'initial_delay_seconds': 5,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 9002,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'grpc',\n 'protocol': 'TCP'},\n {'container_port': 9003,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'grpc-health',\n 'protocol': 'TCP'},\n {'container_port': 9090,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'metrics',\n 'protocol': 'TCP'},\n {'container_port': 5557,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'zmq',\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 3,\n 'grpc': {'port': 9003,\n 'service': 'readiness'},\n 'http_get': None,\n 'initial_delay_seconds': 30,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': None,\n 'requests': {'cpu': '256m',\n 'memory': '500Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n [e2e-llm-inference-service] 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp/tokenizer',\n 'mount_propagation': None,\n 'name': 'tokenizer-uds',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'},\n {'name': 'tls-verification-test-epp-sa-dockercfg-lg5ls'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n [e2e-llm-inference-service] 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'STORAGE_ALLOW_PATTERNS',\n 'value': '["tokenizer.json", '\n '"tokenizer_config.json", '\n '"special_tokens_map.json", '\n '"vocab.json", "merges.txt", '\n '"config.json", '\n '"generation_config.json"]',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n [e2e-llm-inference-service] 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'tls-verification-test-epp-sa',\n 'service_account_name': 'tls-verification-test-epp-sa',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.k [e2e-llm-inference-service] ubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tokenizer-uds',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n [e2e-llm-inference-service] 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tokenizer-tmp',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tokenizer-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_ [e2e-llm-inference-service] map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-nvcq4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n [e2e-llm-inference-service] 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 28, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '256m',\n 'memory': '500Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://06a85f783d00b0ab049cfe211794a4e6e2c9d99ab5b367e8a81dec0d1e949da6',\n 'image': 'ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2',\n 'image_id': 'ghcr.io/llm-d/llm-d-router-endpoint-picker@sha256:06b6c75d77afd0e07053402752a9736c2dfbc12a306d0d37d963aac4c1d4e6a6',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': None,\n 'requests': {'cpu': '256m',\n 'memory': '500Mi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 28, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'D [e2e-llm-inference-service] isabled'},\n {'mount_path': '/tmp/tokenizer',\n 'name': 'tokenizer-uds',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d9461b9905771918ca854c7298a769f287a9c99bed534aa1a989063d165d9e38',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d9461b9905771918ca854c7298a769f287a9c99bed534aa1a989063d165d9e38',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None}, [e2e-llm-inference-service] \n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.65'}],\n 'pod_ip': '10.132.0.65',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': { [e2e-llm-inference-service] },\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n [e2e-llm-inference-service] 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v [e2e-llm-inference-service] 1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n [e2e-llm-inference-service] 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '83807',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n [e2e-llm-inference-service] 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n ' fi\n'\n [e2e-llm-inference-service] '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n ' '\n [e2e-llm-inference-service] 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n [e2e-llm-inference-service] 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n [e2e-llm-inference-service] ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n [e2e-llm-inference-service] 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n [e2e-llm-inference-service] 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n [e2e-llm-inference-service] 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n [e2e-llm-inference-service] 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n [e2e-llm-inference-service] 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None, [e2e-llm-inference-service] \n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n [e2e-llm-inference-service] 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n [e2e-llm-inference-service] 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n [e2e-llm-inference-service] 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e6042236 [e2e-llm-inference-service] 5f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}, {'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'app.kubernetes.io/version': '0.9.0',\n 'certificates.kserve.io/expiration-v2': 'true',\n 'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.65/23"],"mac_address":"0a:58:0a:84:00:41","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_addr [e2e-llm-inference-service] ess":"10.132.0.65/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.65"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:41",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-router-scheduler-7f955879f5-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-router-scheduler',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'pod-template-hash': '7f955879f5'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'.': {},\n 'f:app.kubernetes.io/version': {},\n 'f:certificates.kserve.io/expiration-v2': {}},\n 'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n [e2e-llm-inference-service] 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"b09b6b35-78d8-4429-9188-7f7739e3ad08"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"SSL_CERT_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:grpc': {'.': {},\n 'f:port': {},\n 'f:service': {}},\n 'f:initialDelaySeconds': {},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":5557,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:protocol': {}},\n 'k:{"containerPort":9002,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n 'f:protocol': {}},\n 'k:{"containerPort":9003,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n 'f:protocol': {}},\n 'k:{"containerPort":9090,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:grpc': {'.': {},\n 'f:port': {},\n 'f:service': {}},\n 'f:initialDelaySeconds': {},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n [e2e-llm-inference-service] 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/tmp/tokenizer"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n [e2e-llm-inference-service] 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\ [e2e-llm-inference-service] n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"STORAGE_ALLOW_PATTERNS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n [e2e-llm-inference-service] 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:serviceAccount': {},\n 'f:serviceAccountName': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tokenizer-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tokenizer-tmp"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tokenizer-uds"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal())},\n {'api_version': 'v1',\n [e2e-llm-inference-service] 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.6 [e2e-llm-inference-service] 5"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-router-scheduler-7f955879f5',\n 'uid': 'b09b6b35-78d8-4429-9188-7f7739e3ad08'}],\n 'resource_version': '83144',\n 'self_link': None,\n 'uid': 'a5e7f68b-3c89-43e5-909d-2c8104bb0150'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--config-text',\n 'apiVersion: '\n 'inference.networking.x-k8s.io/v1alpha1\n'\n 'kind: EndpointPickerConfig\n'\n 'plugins:\n'\n '- type: single-profile-handler\n'\n '- type: queue-scorer\n'\n '- type: prefix-cache-scorer\n'\n '- type: max-score-picker\n'\n 'schedulingProfiles:\n'\n '- name: default\n'\n ' plugins:\n'\n ' - pluginRef: queue-scorer\n'\n ' weight: 2\n'\n ' - pluginRef: prefix-cache-scorer\n'\n ' weight: 3\n'\n ' - pluginRef: max-score-picker\n'],\n 'command': ['/app/epp',\n '--pool-name',\n 'tls-verification-test-inference-pool',\n '--pool-namespace',\n 'kserve-ci-e2e-test',\n '--zap-encoder',\n 'json',\n '--grpc-port',\n '9002',\n '--grpc-health-port',\n '9003',\n '--enable-cert-reload=true',\n '--secure-serving=true',\n '--model-server-metrics-scheme=https',\n '--cert-path=/var/run/kserve/tls'],\n 'env': [{'name': 'SSL_CERT_DIR',\n 'value': '/var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n [e2e-llm-inference-service] 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 3,\n 'grpc': {'port': 9003,\n 'service': 'liveness'},\n 'http_get': None,\n 'initial_delay_seconds': 5,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 9002,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'grpc',\n 'protocol': 'TCP'},\n {'container_port': 9003,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'grpc-health',\n 'protocol': 'TCP'},\n {'container_port': 9090,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'metrics',\n 'protocol': 'TCP'},\n {'container_port': 5557,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'zmq',\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 3,\n 'grpc': {'port': 9003,\n 'service': 'readiness'},\n 'http_get': None,\n 'initial_delay_seconds': 30,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': None,\n 'requests': {'cpu': '256m',\n 'memory': '500Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n [e2e-llm-inference-service] 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp/tokenizer',\n 'mount_propagation': None,\n 'name': 'tokenizer-uds',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'},\n {'name': 'tls-verification-test-epp-sa-dockercfg-lg5ls'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n [e2e-llm-inference-service] 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'STORAGE_ALLOW_PATTERNS',\n 'value': '["tokenizer.json", '\n '"tokenizer_config.json", '\n '"special_tokens_map.json", '\n '"vocab.json", "merges.txt", '\n '"config.json", '\n '"generation_config.json"]',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n [e2e-llm-inference-service] 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'tls-verification-test-epp-sa',\n 'service_account_name': 'tls-verification-test-epp-sa',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.ku [e2e-llm-inference-service] bernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tokenizer-uds',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n [e2e-llm-inference-service] 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tokenizer-tmp',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tokenizer-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_m [e2e-llm-inference-service] ap': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-nvcq4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n [e2e-llm-inference-service] 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 28, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '256m',\n 'memory': '500Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://06a85f783d00b0ab049cfe211794a4e6e2c9d99ab5b367e8a81dec0d1e949da6',\n 'image': 'ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2',\n 'image_id': 'ghcr.io/llm-d/llm-d-router-endpoint-picker@sha256:06b6c75d77afd0e07053402752a9736c2dfbc12a306d0d37d963aac4c1d4e6a6',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': None,\n 'requests': {'cpu': '256m',\n 'memory': '500Mi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 28, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Di [e2e-llm-inference-service] sabled'},\n {'mount_path': '/tmp/tokenizer',\n 'name': 'tokenizer-uds',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d9461b9905771918ca854c7298a769f287a9c99bed534aa1a989063d165d9e38',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d9461b9905771918ca854c7298a769f287a9c99bed534aa1a989063d165d9e38',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\ [e2e-llm-inference-service] n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.65'}],\n 'pod_ip': '10.132.0.65',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n [e2e-llm-inference-service] 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n [e2e-llm-inference-service] 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPa [e2e-llm-inference-service] th': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n [e2e-llm-inference-service] 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n [e2e-llm-inference-service] 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n [e2e-llm-inference-service] 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '83807',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n [e2e-llm-inference-service] ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n [e2e-llm-inference-service] ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'echo "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n'\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n [e2e-llm-inference-service] ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n [e2e-llm-inference-service] ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common: '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, an [e2e-llm-inference-service] d UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log- [e2e-llm-inference-service] for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n [e2e-llm-inference-service] {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': ' [e2e-llm-inference-service] 2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n [e2e-llm-inference-service] 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n [e2e-llm-inference-service] 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': [e2e-llm-inference-service] 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n [e2e-llm-inference-service] 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure [e2e-llm-inference-service] _file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n [e2e-llm-inference-service] 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 's [e2e-llm-inference-service] ecret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transit [e2e-llm-inference-service] ion_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n [e2e-llm-inference-service] {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n [e2e-llm-inference-service] 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}},\n {'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'app.kubernetes.io/version': '0.9.0',\n 'certificates.kserve.io/expiration-v2': 'true',\n 'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.65/23"],"mac_address":"0a:58:0a:84:00:41","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.65/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.65"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:41",\n'\n [e2e-llm-inference-service] ' '\n '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': None,\n 'deletion_timestamp': None,\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-router-scheduler-7f955879f5-',\n 'generation': 1,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-router-scheduler',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'pod-template-hash': '7f955879f5'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'.': {},\n 'f:app.kubernetes.io/version': {},\n 'f:certificates.kserve.io/expiration-v2': {}},\n 'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"b09b6b35-78d8-4429-9188-7f7739e3ad08"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:args': {},\n 'f:command': {},\n [e2e-llm-inference-service] 'f:env': {'.': {},\n 'k:{"name":"SSL_CERT_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:grpc': {'.': {},\n 'f:port': {},\n 'f:service': {}},\n 'f:initialDelaySeconds': {},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":5557,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n 'f:protocol': {}},\n 'k:{"containerPort":9002,"protocol":"TCP"}': {'.': {},\n [e2e-llm-inference-service] 'f:containerPort': {},\n 'f:name': {},\n 'f:protocol': {}},\n 'k:{"containerPort":9003,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n 'f:protocol': {}},\n 'k:{"containerPort":9090,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:name': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:grpc': {'.': {},\n 'f:port': {},\n 'f:service': {}},\n 'f:initialDelaySeconds': {},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n [e2e-llm-inference-service] 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/tmp/tokenizer"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:val [e2e-llm-inference-service] ueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"STORAGE_ALLOW_PATTERNS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n [e2e-llm-inference-service] 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:serviceAccount': {},\n 'f:serviceAccountName': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tokenizer-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tokenizer-tmp"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tokenizer-uds"}': {'.': {},\n 'f:emptyDir': {},\n [e2e-llm-inference-service] 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n [e2e-llm-inference-service] 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.65"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-router-scheduler-7f955879f5',\n 'uid': 'b09b6b35-78d8-4429-9188-7f7739e3ad08'}],\n 'resource_version': '83144',\n 'self_link': None,\n 'uid': 'a5e7f68b-3c89-43e5-909d-2c8104bb0150'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': ['--config-text',\n 'apiVersion: '\n 'inference.networking.x-k8s.io/v1alpha1\n'\n 'kind: EndpointPickerConfig\n'\n 'plugins:\n'\n '- type: single-profile-handler\n'\n '- type: queue-scorer\n'\n '- type: prefix-cache-scorer\n'\n '- type: max-score-picker\n'\n 'schedulingProfiles:\n'\n '- name: default\n'\n ' plugins:\n'\n ' - pluginRef: queue-scorer\n'\n ' weight: 2\n'\n ' - p [e2e-llm-inference-service] luginRef: '\n 'prefix-cache-scorer\n'\n ' weight: 3\n'\n ' - pluginRef: '\n 'max-score-picker\n'],\n 'command': ['/app/epp',\n '--pool-name',\n 'tls-verification-test-inference-pool',\n '--pool-namespace',\n 'kserve-ci-e2e-test',\n '--zap-encoder',\n 'json',\n '--grpc-port',\n '9002',\n '--grpc-health-port',\n '9003',\n '--enable-cert-reload=true',\n '--secure-serving=true',\n '--model-server-metrics-scheme=https',\n '--cert-path=/var/run/kserve/tls'],\n 'env': [{'name': 'SSL_CERT_DIR',\n 'value': '/var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 3,\n 'grpc': {'port': 9003,\n 'service': 'liveness'},\n 'http_get': None,\n 'initial_delay_seconds': 5,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 9002,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'grpc',\n 'protocol': 'TCP'},\n {'container_port': 9003,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'grpc-health',\n 'protocol': 'TCP'},\n {'container_port': [e2e-llm-inference-service] 9090,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'metrics',\n 'protocol': 'TCP'},\n {'container_port': 5557,\n 'host_ip': None,\n 'host_port': None,\n 'name': 'zmq',\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 3,\n 'grpc': {'port': 9003,\n 'service': 'readiness'},\n 'http_get': None,\n 'initial_delay_seconds': 30,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': None,\n 'requests': {'cpu': '256m',\n 'memory': '500Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n [e2e-llm-inference-service] 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp/tokenizer',\n 'mount_propagation': None,\n 'name': 'tokenizer-uds',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'},\n {'name': 'tls-verification-test-epp-sa-dockercfg-lg5ls'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n [e2e-llm-inference-service] 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None},\n {'name': 'STORAGE_ALLOW_PATTERNS',\n 'value': '["tokenizer.json", '\n '"tokenizer_config.json", '\n '"special_tokens_map.json", '\n '"vocab.json", '\n '"merges.txt", '\n '"config.json", '\n '"generation_config.json"]',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL [e2e-llm-inference-service] ']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n [e2e-llm-inference-service] 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'tls-verification-test-epp-sa',\n 'service_account_name': 'tls-verification-test-epp-sa',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n [e2e-llm-inference-service] 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tokenizer-uds',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tokenizer-tmp',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': Non [e2e-llm-inference-service] e,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tokenizer-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-nvcq4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n [e2e-llm-inference-service] 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason' [e2e-llm-inference-service] : None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 28, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 12, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '256m',\n 'memory': '500Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://06a85f783d00b0ab049cfe211794a4e6e2c9d99ab5b367e8a81dec0d1e949da6',\n 'image': 'ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2',\n 'image_id': 'ghcr.io/llm-d/llm-d-router-endpoint-picker@sha256:06b6c75d77afd0e07053402752a9736c2dfbc12a306d0d37d963aac4c1d4e6a6',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': None,\n 'requests': {'cpu': '256m',\n 'memory': '500Mi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 28, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume [e2e-llm-inference-service] _mounts': [{'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/tmp/tokenizer',\n 'name': 'tokenizer-uds',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://d9461b9905771918ca854c7298a769f287a9c99bed534aa1a989063d165d9e38',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://d9461b9905771918ca854c7298a769f287a9c99bed534aa1a989063d165d9e38',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal())},\n [e2e-llm-inference-service] 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-nvcq4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.65'}],\n 'pod_ip': '10.132.0.65',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '91196',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for tls-verification-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 29, 34, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n ' [e2e-llm-inference-service] fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContai [e2e-llm-inference-service] nerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 28, 34, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '91231',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n [e2e-llm-inference-service] ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n [e2e-llm-inference-service] ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n [e2e-llm-inference-service] ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n [e2e-llm-inference-service] ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n [e2e-llm-inference-service] ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n [e2e-llm-inference-service] 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n [e2e-llm-inference-service] 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-c [e2e-llm-inference-service] reds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_moun [e2e-llm-inference-service] t': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n [e2e-llm-inference-service] 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_v [e2e-llm-inference-service] olume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n [e2e-llm-inference-service] 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mo [e2e-llm-inference-service] de': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n [e2e-llm-inference-service] 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sh [e2e-llm-inference-service] a256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 29, 34, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'f [e2e-llm-inference-service] ields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContain [e2e-llm-inference-service] erStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 28, 34, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '91231',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n [e2e-llm-inference-service] ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n [e2e-llm-inference-service] ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n [e2e-llm-inference-service] ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n [e2e-llm-inference-service] ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n [e2e-llm-inference-service] ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n [e2e-llm-inference-service] 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n [e2e-llm-inference-service] 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-cr [e2e-llm-inference-service] eds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount [e2e-llm-inference-service] ': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n [e2e-llm-inference-service] 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_vo [e2e-llm-inference-service] lume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n [e2e-llm-inference-service] 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mod [e2e-llm-inference-service] e': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n [e2e-llm-inference-service] 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha [e2e-llm-inference-service] 256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"1 [e2e-llm-inference-service] 0.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 29, 34, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n [e2e-llm-inference-service] 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n [e2e-llm-inference-service] 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n [e2e-llm-inference-service] 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n [e2e-llm-inference-service] 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n [e2e-llm-inference-service] 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.' [e2e-llm-inference-service] : {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 28, 34, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '91231',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1 [e2e-llm-inference-service] /state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'echo "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST: [e2e-llm-inference-service] -${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n'\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n [e2e-llm-inference-service] ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common [e2e-llm-inference-service] : '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n [e2e-llm-inference-service] 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n [e2e-llm-inference-service] 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'f [e2e-llm-inference-service] ailure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n [e2e-llm-inference-service] 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n [e2e-llm-inference-service] {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n [e2e-llm-inference-service] {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': Tr [e2e-llm-inference-service] ue,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n [e2e-llm-inference-service] 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n [e2e-llm-inference-service] 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'project [e2e-llm-inference-service] ed': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n [e2e-llm-inference-service] {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n [e2e-llm-inference-service] 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n [e2e-llm-inference-service] 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n [e2e-llm-inference-service] 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}], [e2e-llm-inference-service] \n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '91345',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for tls-verification-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 29, 34, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n ' [e2e-llm-inference-service] fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContai [e2e-llm-inference-service] nerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 28, 34, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '91231',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n [e2e-llm-inference-service] ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n [e2e-llm-inference-service] ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n [e2e-llm-inference-service] ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n [e2e-llm-inference-service] ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n [e2e-llm-inference-service] ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n [e2e-llm-inference-service] 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n [e2e-llm-inference-service] 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-c [e2e-llm-inference-service] reds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_moun [e2e-llm-inference-service] t': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n [e2e-llm-inference-service] 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_v [e2e-llm-inference-service] olume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n [e2e-llm-inference-service] 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mo [e2e-llm-inference-service] de': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n [e2e-llm-inference-service] 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sh [e2e-llm-inference-service] a256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 29, 34, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'f [e2e-llm-inference-service] ields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContain [e2e-llm-inference-service] erStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 28, 34, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '91231',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n [e2e-llm-inference-service] ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n [e2e-llm-inference-service] ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n [e2e-llm-inference-service] ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n [e2e-llm-inference-service] ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n [e2e-llm-inference-service] ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n [e2e-llm-inference-service] 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n [e2e-llm-inference-service] 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-cr [e2e-llm-inference-service] eds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount [e2e-llm-inference-service] ': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n [e2e-llm-inference-service] 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_vo [e2e-llm-inference-service] lume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n [e2e-llm-inference-service] 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mod [e2e-llm-inference-service] e': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n [e2e-llm-inference-service] 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha [e2e-llm-inference-service] 256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"1 [e2e-llm-inference-service] 0.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 29, 34, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n [e2e-llm-inference-service] 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n [e2e-llm-inference-service] 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n [e2e-llm-inference-service] 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n [e2e-llm-inference-service] 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n [e2e-llm-inference-service] 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.' [e2e-llm-inference-service] : {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 28, 34, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '91231',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1 [e2e-llm-inference-service] /state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'echo "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST: [e2e-llm-inference-service] -${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n'\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n [e2e-llm-inference-service] ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common [e2e-llm-inference-service] : '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n [e2e-llm-inference-service] 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n [e2e-llm-inference-service] 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'f [e2e-llm-inference-service] ailure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n [e2e-llm-inference-service] 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n [e2e-llm-inference-service] {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n [e2e-llm-inference-service] {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': Tr [e2e-llm-inference-service] ue,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n [e2e-llm-inference-service] 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n [e2e-llm-inference-service] 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'project [e2e-llm-inference-service] ed': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n [e2e-llm-inference-service] {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n [e2e-llm-inference-service] 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n [e2e-llm-inference-service] 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n [e2e-llm-inference-service] 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}], [e2e-llm-inference-service] \n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '91381',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: 1 pod(s) for tls-verification-test still terminating [e2e-llm-inference-service] assert not [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 29, 34, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n ' [e2e-llm-inference-service] fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContai [e2e-llm-inference-service] nerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 28, 34, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '91231',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n [e2e-llm-inference-service] ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n [e2e-llm-inference-service] ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n [e2e-llm-inference-service] ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n [e2e-llm-inference-service] ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n [e2e-llm-inference-service] ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n [e2e-llm-inference-service] 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n [e2e-llm-inference-service] 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-c [e2e-llm-inference-service] reds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_moun [e2e-llm-inference-service] t': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n [e2e-llm-inference-service] 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_v [e2e-llm-inference-service] olume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n [e2e-llm-inference-service] 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mo [e2e-llm-inference-service] de': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n [e2e-llm-inference-service] 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sh [e2e-llm-inference-service] a256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}] [e2e-llm-inference-service] + where [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"10.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' "ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' ],\n'\n ' "mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' "dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 29, 34, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n [e2e-llm-inference-service] 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n [e2e-llm-inference-service] 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n [e2e-llm-inference-service] 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n [e2e-llm-inference-service] 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n [e2e-llm-inference-service] 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'f [e2e-llm-inference-service] ields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContain [e2e-llm-inference-service] erStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 28, 34, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '91231',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f /etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = "true" ]; '\n 'then\n'\n ' echo "Trying to infer RoCE configs '\n '... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat /proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] Discovering active '\n 'HCAs ..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 devices found '\n 'in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; do\n'\n " # Ensure it's a directory before "\n 'proceeding\n'\n ' if [ -d "$hca_dir" ]; then\n'\n ' hca_name=$(basename '\n '"$hca_dir")\n'\n [e2e-llm-inference-service] ' '\n 'port_state_file="$hca_dir/ports/1/state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] Check if '\n 'the port state file ${port_state_file} '\n 'exists and contains \'ACTIVE\'"\n'\n ' if [ -f "$port_state_file" ] '\n '&& grep -q "ACTIVE" "$port_state_file" '\n '&& grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; then\n'\n ' echo "[Infer RoCE] Found '\n 'active HCA: $hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'Skipping inactive or down HCA: '\n '$hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any active HCAs\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' # Join the array elements with a '\n 'comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in "${active_hcas[@]}"; '\n 'do\n'\n ' hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' active_hca_list=$(IFS=,; echo '\n '"${active_hcas[*]}")\n'\n ' hca_port_pairs_list=$(IFS=,; echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] Setting active '\n 'HCAs: ${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST:-${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] WARNING: No '\n 'active RoCE HCAs found. NCCL_IB_HCA '\n 'will not be set."\n'\n [e2e-llm-inference-service] ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} -gt 0 ]; '\n 'then\n'\n ' echo "[Infer RoCE] Finding '\n 'GID_INDEX for each active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV environments, find '\n 'the most common IPv4 RoCE v2 GID index '\n 'across all HCAs\n'\n ' declare -A gid_index_count\n'\n ' declare -A hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] Processing '\n 'HCA: ${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 IPv4 GIDs '\n 'for this HCA and count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' idx=$(basename '\n '"$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check for IPv4 GID '\n '(contains ffff:)\n'\n ' if [ -f "$gid_file" ] '\n '&& grep -q "ffff:" "$gid_file"; then\n'\n ' gid_value=$(cat '\n '"$gid_file" 2>/dev/null || echo "")\n'\n ' echo "[Infer '\n 'RoCE] Found IPv4 RoCE v2 GID for '\n '${hca_name}: index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break # Use '\n 'first found IPv4 GID per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common GID index '\n '(most likely to be consistent across '\n 'nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; do\n'\n [e2e-llm-inference-service] ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] GID_INDEX '\n '${idx} found on ${count} HCAs"\n'\n ' if [ $count -gt $max_count ]; '\n 'then\n'\n ' max_count=$count\n'\n ' best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic fallback if '\n 'tied - prefer index 3 (SR-IOV '\n 'standard)\n'\n ' if [ ${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] Multiple '\n 'GID indices found, selecting most '\n 'common: ${best_gid_index}"\n'\n " # If there's a tie, prefer "\n "index 3 as it's most common in SR-IOV "\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" ] && [ '\n '"${gid_index_count[\'3\']}" -eq '\n '"$max_count" ]; then\n'\n ' best_gid_index="3"\n'\n ' echo "[Infer RoCE] Using '\n 'deterministic fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX is already '\n 'set via environment variables\n'\n ' if [ -n "${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] Using '\n 'pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} for '\n 'NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n "$best_gid_index" ]; '\n 'then\n'\n ' echo "[Infer RoCE] Selected '\n 'GID_INDEX: ${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n [e2e-llm-inference-service] ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] Exported '\n 'GID_INDEX=${best_gid_index} for NCCL, '\n 'NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] ERROR: No '\n 'valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No active HCAs '\n 'found, skipping GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# --disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm --version '\n "2>/dev/null | tail -1 | awk '{print "\n "$NF}')\n"\n 'echo "[access-log-detect] vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.16.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed in vLLM '\n '0.18.0 (vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ "$(printf '\n '\'%s\\n%s\\n\' "0.18.0" '\n '"${VLLM_VERSION}" | sort -V | head -1)" '\n '= "0.18.0" ]; then\n'\n ' '\n 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve /mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n [e2e-llm-inference-service] ' ${SHUTDOWN_TIMEOUT_ARGS} \\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt \\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key \\\n'\n ' ${VLLM_ADDITIONAL_ARGS} \\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'failure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n [e2e-llm-inference-service] 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2', 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n [e2e-llm-inference-service] 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-cr [e2e-llm-inference-service] eds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount [e2e-llm-inference-service] ': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n [e2e-llm-inference-service] 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory', 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_vo [e2e-llm-inference-service] lume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n [e2e-llm-inference-service] 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None, 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mod [e2e-llm-inference-service] e': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n [e2e-llm-inference-service] 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha [e2e-llm-inference-service] 256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}] = {'api_version': 'v1',\n 'items': [{'api_version': None,\n 'kind': None,\n 'metadata': {'annotations': {'k8s.ovn.org/pod-networks': '{"default":{"ip_addresses":["10.132.0.64/23"],"mac_address":"0a:58:0a:84:00:40","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.64/23","gateway_ip":"1 [e2e-llm-inference-service] 0.132.0.1","role":"primary"}}',\n 'k8s.v1.cni.cncf.io/network-status': '[{\n'\n ' '\n '"name": '\n '"ovn-kubernetes",\n'\n ' '\n '"interface": '\n '"eth0",\n'\n ' '\n '"ips": '\n '[\n'\n ' '\n '"10.132.0.64"\n'\n ' '\n '],\n'\n ' '\n '"mac": '\n '"0a:58:0a:84:00:40",\n'\n ' '\n '"default": '\n 'true,\n'\n ' '\n '"dns": '\n '{}\n'\n '}]',\n 'openshift.io/scc': 'restricted-v2',\n 'seccomp.security.alpha.kubernetes.io/pod': 'runtime/default',\n 'security.openshift.io/validated-scc-subject-type': 'user'},\n 'creation_timestamp': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'deletion_grace_period_seconds': 60,\n 'deletion_timestamp': datetime.datetime(2026, 7, 2, 14, 29, 34, tzinfo=tzlocal()),\n 'finalizers': None,\n 'generate_name': 'tls-verification-test-kserve-6d6946dcd6-',\n 'generation': 2,\n 'labels': {'app.kubernetes.io/component': 'llminferenceservice-workload',\n 'app.kubernetes.io/name': 'tls-verification-test',\n 'app.kubernetes.io/part-of': 'llminferenceservice',\n 'kserve.io/component': 'workload',\n 'llm-d.ai/role': 'both',\n 'pod-template-hash': '6d6946dcd6'},\n 'managed_fields': [{'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.ovn.org/pod-networks': {}}}},\n 'manager': 'ip-10-0-134-1',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n [e2e-llm-inference-service] 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:generateName': {},\n 'f:labels': {'.': {},\n 'f:app.kubernetes.io/component': {},\n 'f:app.kubernetes.io/name': {},\n 'f:app.kubernetes.io/part-of': {},\n 'f:kserve.io/component': {},\n 'f:llm-d.ai/role': {},\n 'f:pod-template-hash': {}},\n 'f:ownerReferences': {'.': {},\n 'k:{"uid":"617afe42-e508-4967-bf73-581781225c99"}': {}}},\n 'f:spec': {'f:containers': {'k:{"name":"main"}': {'.': {},\n 'f:command': {},\n 'f:env': {'.': {},\n 'k:{"name":"HF_HUB_CACHE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HOME"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"TORCHINDUCTOR_CACHE_DIR"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"USER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_CPU_KVCACHE_SPACE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"VLLM_ENABLE_V1_MULTIPROCESSING"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"VLLM_LOGGING_LEVEL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:lifecycle': {'.': {},\n 'f:preStop': {'.': {},\n 'f:exec': {'.': {},\n 'f:command': {}}}},\n 'f:livenessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:name': {},\n 'f:ports': {'.': {},\n 'k:{"containerPort":8000,"protocol":"TCP"}': {'.': {},\n 'f:containerPort': {},\n 'f:protocol': {}}},\n [e2e-llm-inference-service] 'f:readinessProbe': {'.': {},\n 'f:failureThreshold': {},\n 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:securityContext': {'.': {},\n 'f:allowPrivilegeEscalation': {},\n 'f:capabilities': {'.': {},\n 'f:drop': {}},\n 'f:readOnlyRootFilesystem': {},\n 'f:runAsNonRoot': {},\n 'f:seccompProfile': {'.': {},\n 'f:type': {}}},\n 'f:startupProbe': {'.': {},\n 'f:failureThreshold': {},\n [e2e-llm-inference-service] 'f:httpGet': {'.': {},\n 'f:path': {},\n 'f:port': {},\n 'f:scheme': {}},\n 'f:periodSeconds': {},\n 'f:successThreshold': {},\n 'f:timeoutSeconds': {}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/dev/shm"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/home"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}},\n 'k:{"mountPath":"/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}},\n 'k:{"mountPath":"/tmp"}': {'.': {},\n 'f:mountPath': {},\n [e2e-llm-inference-service] 'f:name': {}},\n 'k:{"mountPath":"/var/run/kserve/tls"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {},\n 'f:readOnly': {}}}}},\n 'f:dnsPolicy': {},\n 'f:enableServiceLinks': {},\n 'f:initContainers': {'.': {},\n 'k:{"name":"storage-initializer"}': {'.': {},\n 'f:args': {},\n 'f:env': {'.': {},\n 'k:{"name":"AWS_ACCESS_KEY_ID"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"AWS_CA_BUNDLE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_CA_BUNDLE_CONFIGMAP"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"AWS_ENDPOINT_URL"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n [e2e-llm-inference-service] 'k:{"name":"AWS_SECRET_ACCESS_KEY"}': {'.': {},\n 'f:name': {},\n 'f:valueFrom': {'.': {},\n 'f:secretKeyRef': {}}},\n 'k:{"name":"HF_HUB_ENABLE_HF_TRANSFER"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_HIGH_PERFORMANCE"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"HF_XET_NUM_CONCURRENT_RANGE_GETS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_ENDPOINT"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_USE_HTTPS"}': {'.': {},\n 'f:name': {},\n 'f:value': {}},\n 'k:{"name":"S3_VERIFY_SSL"}': {'.': {},\n 'f:name': {},\n [e2e-llm-inference-service] 'f:value': {}}},\n 'f:image': {},\n 'f:imagePullPolicy': {},\n 'f:name': {},\n 'f:resources': {'.': {},\n 'f:limits': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}},\n 'f:requests': {'.': {},\n 'f:cpu': {},\n 'f:memory': {}}},\n 'f:terminationMessagePath': {},\n 'f:terminationMessagePolicy': {},\n 'f:volumeMounts': {'.': {},\n 'k:{"mountPath":"/mnt/models"}': {'.': {},\n 'f:mountPath': {},\n 'f:name': {}}}}},\n 'f:restartPolicy': {},\n 'f:schedulerName': {},\n 'f:securityContext': {},\n 'f:terminationGracePeriodSeconds': {},\n 'f:volumes': {'.': {},\n 'k:{"name":"dshm"}': {'.': {},\n 'f:emptyDir': {'.': {},\n 'f:medium': {},\n 'f:sizeLimit': {}},\n 'f:name': {}},\n 'k:{"name":"home"}': {'.': {},\n [e2e-llm-inference-service] 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"kserve-provision-location"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"model-cache"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}},\n 'k:{"name":"tls-certs"}': {'.': {},\n 'f:name': {},\n 'f:secret': {'.': {},\n 'f:defaultMode': {},\n 'f:secretName': {}}},\n 'k:{"name":"tmp-dir"}': {'.': {},\n 'f:emptyDir': {},\n 'f:name': {}}}}},\n 'manager': 'kube-controller-manager',\n 'operation': 'Update',\n 'subresource': None,\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:metadata': {'f:annotations': {'f:k8s.v1.cni.cncf.io/network-status': {}}}},\n 'manager': 'multus-daemon',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n {'api_version': 'v1',\n 'fields_type': 'FieldsV1',\n 'fields_v1': {'f:status': {'f:conditions': {'k:{"type":"ContainersReady"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"Initialized"}': {'.' [e2e-llm-inference-service] : {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodReadyToStartContainers"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}},\n 'k:{"type":"PodScheduled"}': {'f:observedGeneration': {}},\n 'k:{"type":"Ready"}': {'.': {},\n 'f:lastProbeTime': {},\n 'f:lastTransitionTime': {},\n 'f:observedGeneration': {},\n 'f:status': {},\n 'f:type': {}}},\n 'f:containerStatuses': {},\n 'f:hostIP': {},\n 'f:hostIPs': {},\n 'f:initContainerStatuses': {},\n 'f:observedGeneration': {},\n 'f:phase': {},\n 'f:podIP': {},\n 'f:podIPs': {'.': {},\n 'k:{"ip":"10.132.0.64"}': {'.': {},\n 'f:ip': {}}},\n 'f:startTime': {}}},\n 'manager': 'kubelet',\n 'operation': 'Update',\n 'subresource': 'status',\n 'time': datetime.datetime(2026, 7, 2, 14, 28, 34, tzinfo=tzlocal())}],\n 'name': 'tls-verification-test-kserve-6d6946dcd6-bdddw',\n [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test',\n 'owner_references': [{'api_version': 'apps/v1',\n 'block_owner_deletion': True,\n 'controller': True,\n 'kind': 'ReplicaSet',\n 'name': 'tls-verification-test-kserve-6d6946dcd6',\n 'uid': '617afe42-e508-4967-bf73-581781225c99'}],\n 'resource_version': '91231',\n 'self_link': None,\n 'uid': 'dc7326bd-f917-466c-922f-739706ba0fda'},\n 'spec': {'active_deadline_seconds': None,\n 'affinity': None,\n 'automount_service_account_token': None,\n 'containers': [{'args': None,\n 'command': ['/bin/bash',\n '-c',\n 'if [ -f '\n '/etc/profile.d/ibm-aiu-setup.sh '\n ']; then\n'\n ' source '\n '/etc/profile.d/ibm-aiu-setup.sh\n'\n 'fi\n'\n '\n'\n 'if [ "$KSERVE_INFER_ROCE" = '\n '"true" ]; then\n'\n ' echo "Trying to infer RoCE '\n 'configs ... "\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gids/* '\n '2>/dev/null\n'\n ' grep -H . '\n '/sys/class/infiniband/*/ports/*/gid_attrs/types/* '\n '2>/dev/null\n'\n '\n'\n ' cat '\n '/proc/driver/nvidia/params\n'\n '\n'\n ' '\n 'KSERVE_INFER_IB_GID_INDEX_GREP=${KSERVE_INFER_IB_GID_INDEX_GREP:-"RoCE '\n 'v2"}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Discovering active HCAs '\n '..."\n'\n ' active_hcas=()\n'\n ' # Loop through all mlx5 '\n 'devices found in sysfs\n'\n ' for hca_dir in '\n '/sys/class/infiniband/mlx5_*; '\n 'do\n'\n " # Ensure it's a "\n 'directory before proceeding\n'\n ' if [ -d "$hca_dir" ]; '\n 'then\n'\n ' '\n 'hca_name=$(basename '\n '"$hca_dir")\n'\n ' '\n 'port_state_file="$hca_dir/ports/1 [e2e-llm-inference-service] /state" '\n '# Assume port 1\n'\n ' '\n 'type_file="$hca_dir/ports/1/gid_attrs/types/*"\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Check if the port state file '\n '${port_state_file} exists '\n 'and contains \'ACTIVE\'"\n'\n ' if [ -f '\n '"$port_state_file" ] && grep '\n '-q "ACTIVE" '\n '"$port_state_file" && grep '\n '-q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '${type_file} 2>/dev/null; '\n 'then\n'\n ' echo "[Infer '\n 'RoCE] Found active HCA: '\n '$hca_name"\n'\n ' '\n 'active_hcas+=("$hca_name")\n'\n ' else\n'\n ' echo "[Infer '\n 'RoCE] Skipping inactive or '\n 'down HCA: $hca_name"\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Check if we found any '\n 'active HCAs\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' # Join the array '\n 'elements with a comma\n'\n ' hca_port_pairs=()\n'\n ' for hca in '\n '"${active_hcas[@]}"; do\n'\n ' '\n 'hca_port_pairs+=("${hca}:1")\n'\n ' done\n'\n '\n'\n ' '\n 'active_hca_list=$(IFS=,; '\n 'echo "${active_hcas[*]}")\n'\n ' '\n 'hca_port_pairs_list=$(IFS=,; '\n 'echo '\n '"${hca_port_pairs[*]}")\n'\n ' echo "[Infer RoCE] '\n 'Setting active HCAs: '\n '${active_hca_list}"\n'\n ' export '\n 'NCCL_IB_HCA=${NCCL_IB_HCA:-${active_hca_list}}\n'\n ' export '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST: [e2e-llm-inference-service] -${hca_port_pairs_list}}\n'\n ' export '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES:-${hca_port_pairs_list}}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'NCCL_IB_HCA=${NCCL_IB_HCA}"\n'\n ' echo "[Infer RoCE] '\n 'NVSHMEM_HCA_LIST=${NVSHMEM_HCA_LIST}"\n'\n ' echo "[Infer RoCE] '\n 'UCX_NET_DEVICES=${UCX_NET_DEVICES}"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'WARNING: No active RoCE HCAs '\n 'found. NCCL_IB_HCA will not '\n 'be set."\n'\n ' fi\n'\n '\n'\n ' if [ ${#active_hcas[@]} '\n '-gt 0 ]; then\n'\n ' echo "[Infer RoCE] '\n 'Finding GID_INDEX for each '\n 'active HCA (SR-IOV '\n 'compatible)..."\n'\n '\n'\n ' # For SR-IOV '\n 'environments, find the most '\n 'common IPv4 RoCE v2 GID '\n 'index across all HCAs\n'\n ' declare -A '\n 'gid_index_count\n'\n ' declare -A '\n 'hca_gid_index\n'\n '\n'\n ' for hca_name in '\n '"${active_hcas[@]}"; do\n'\n ' echo "[Infer RoCE] '\n 'Processing HCA: '\n '${hca_name}"\n'\n '\n'\n ' # Find all RoCE v2 '\n 'IPv4 GIDs for this HCA and '\n 'count by index\n'\n ' for tpath in '\n '/sys/class/infiniband/${hca_name}/ports/1/gid_attrs/types/*; '\n 'do\n'\n ' if grep -q '\n '"${KSERVE_INFER_IB_GID_INDEX_GREP}" '\n '"$tpath" 2>/dev/null; then\n'\n ' '\n 'idx=$(basename "$tpath")\n'\n ' '\n 'gid_file="/sys/class/infiniband/${hca_name}/ports/1/gids/${idx}"\n'\n ' # Check '\n 'for IPv4 GID (contains '\n 'ffff:)\n'\n [e2e-llm-inference-service] ' if [ -f '\n '"$gid_file" ] && grep -q '\n '"ffff:" "$gid_file"; then\n'\n ' '\n 'gid_value=$(cat "$gid_file" '\n '2>/dev/null || echo "")\n'\n ' echo '\n '"[Infer RoCE] Found IPv4 '\n 'RoCE v2 GID for ${hca_name}: '\n 'index=${idx}, '\n 'gid=${gid_value}"\n'\n ' '\n 'hca_gid_index["${hca_name}"]="${idx}"\n'\n ' '\n 'gid_index_count["${idx}"]=$((${gid_index_count["${idx}"]} '\n '+ 1))\n'\n ' break '\n '# Use first found IPv4 GID '\n 'per HCA\n'\n ' fi\n'\n ' fi\n'\n ' done\n'\n ' done\n'\n '\n'\n ' # Find the most common '\n 'GID index (most likely to be '\n 'consistent across nodes)\n'\n ' best_gid_index=""\n'\n ' max_count=0\n'\n ' for idx in '\n '"${!gid_index_count[@]}"; '\n 'do\n'\n ' '\n 'count=${gid_index_count["${idx}"]}\n'\n ' echo "[Infer RoCE] '\n 'GID_INDEX ${idx} found on '\n '${count} HCAs"\n'\n ' if [ $count -gt '\n '$max_count ]; then\n'\n ' '\n 'max_count=$count\n'\n ' '\n 'best_gid_index="$idx"\n'\n ' fi\n'\n ' done\n'\n '\n'\n ' # Use deterministic '\n 'fallback if tied - prefer '\n 'index 3 (SR-IOV standard)\n'\n ' if [ '\n '${#gid_index_count[@]} -gt 1 '\n ']; then\n'\n ' echo "[Infer RoCE] '\n 'Multiple GID indices found, '\n 'selecting most common [e2e-llm-inference-service] : '\n '${best_gid_index}"\n'\n " # If there's a "\n "tie, prefer index 3 as it's "\n 'most common in SR-IOV '\n 'setups\n'\n ' if [ -n '\n '"${gid_index_count[\'3\']}" '\n '] && [ '\n '"${gid_index_count[\'3\']}" '\n '-eq "$max_count" ]; then\n'\n ' '\n 'best_gid_index="3"\n'\n ' echo "[Infer '\n 'RoCE] Using deterministic '\n 'fallback: GID_INDEX=3 '\n '(SR-IOV standard)"\n'\n ' fi\n'\n ' fi\n'\n '\n'\n ' # Check if GID_INDEX '\n 'is already set via '\n 'environment variables\n'\n ' if [ -n '\n '"${NCCL_IB_GID_INDEX}" ]; '\n 'then\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'from environment"\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$NCCL_IB_GID_INDEX}\n'\n ' echo "[Infer RoCE] '\n 'Using pre-configured '\n 'GID_INDEX=${NCCL_IB_GID_INDEX} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' elif [ -n '\n '"$best_gid_index" ]; then\n'\n ' echo "[Infer RoCE] '\n 'Selected GID_INDEX: '\n '${best_gid_index} (found on '\n '${max_count} HCAs)"\n'\n '\n'\n ' export '\n 'NCCL_IB_GID_INDEX=${NCCL_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'NVSHMEM_IB_GID_INDEX=${NVSHMEM_IB_GID_INDEX:-$best_gid_index}\n'\n ' export '\n 'UCX_IB_GID_INDEX=${UCX_IB_GID_INDEX:-$best_gid_index}\n'\n '\n'\n ' echo "[Infer RoCE] '\n 'Exported '\n [e2e-llm-inference-service] 'GID_INDEX=${best_gid_index} '\n 'for NCCL, NVSHMEM, and UCX"\n'\n ' else\n'\n ' echo "[Infer RoCE] '\n 'ERROR: No valid IPv4 '\n '${KSERVE_INFER_IB_GID_INDEX_GREP} '\n 'GID_INDEX found on any '\n 'HCA."\n'\n ' fi\n'\n ' else\n'\n ' echo "[Infer RoCE] No '\n 'active HCAs found, skipping '\n 'GID_INDEX inference."\n'\n ' fi\n'\n 'fi\n'\n '\n'\n '# '\n '--disable-access-log-for-endpoints '\n 'landed in vLLM 0.16.0 '\n '(vllm-project/vllm#30011).\n'\n '# Older versions still need '\n 'the blanket '\n '--disable-uvicorn-access-log.\n'\n 'ACCESS_LOG_ARGS="--disable-uvicorn-access-log"\n'\n 'VLLM_VERSION=$(vllm '\n '--version 2>/dev/null | tail '\n "-1 | awk '{print $NF}')\n"\n 'echo "[access-log-detect] '\n 'vllm '\n 'version=\'${VLLM_VERSION}\'"\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.16.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.16.0" ]; then\n'\n ' '\n 'ACCESS_LOG_ARGS="--disable-access-log-for-endpoints '\n '/health,/metrics,/ping"\n'\n 'fi\n'\n 'echo "[access-log-detect] '\n 'selected '\n 'ACCESS_LOG_ARGS=\'${ACCESS_LOG_ARGS}\'"\n'\n '\n'\n '# --shutdown-timeout landed '\n 'in vLLM 0.18.0 '\n '(vllm-project/vllm#36666).\n'\n 'SHUTDOWN_TIMEOUT_ARGS=""\n'\n 'if [[ "$VLLM_VERSION" =~ '\n '^[0-9]+\\.[0-9]+ ]] && [ '\n '"$(printf \'%s\\n%s\\n\' '\n '"0.18.0" "${VLLM_VERSION}" | '\n 'sort -V | head -1)" = '\n '"0.18.0" ]; then\n'\n ' '\n [e2e-llm-inference-service] 'SHUTDOWN_TIMEOUT_ARGS="--shutdown-timeout '\n '40"\n'\n 'fi\n'\n '\n'\n 'eval "exec vllm serve '\n '/mnt/models \\\n'\n ' --served-model-name '\n '"facebook/opt-125m" '\n '"publishers/kserve-ci-e2e-test/models/facebook/opt-125m" '\n '\\\n'\n ' --port 8000 \\\n'\n ' ${ACCESS_LOG_ARGS} \\\n'\n ' ${SHUTDOWN_TIMEOUT_ARGS} '\n '\\\n'\n ' --enable-ssl-refresh \\\n'\n ' --ssl-certfile '\n '/var/run/kserve/tls/tls.crt '\n '\\\n'\n ' --ssl-keyfile '\n '/var/run/kserve/tls/tls.key '\n '\\\n'\n ' ${VLLM_ADDITIONAL_ARGS} '\n '\\\n'\n ' $@"',\n '--'],\n 'env': [{'name': 'HOME',\n 'value': '/home',\n 'value_from': None},\n {'name': 'VLLM_LOGGING_LEVEL',\n 'value': 'DEBUG',\n 'value_from': None},\n {'name': 'VLLM_CPU_KVCACHE_SPACE',\n 'value': '1',\n 'value_from': None},\n {'name': 'VLLM_ENABLE_V1_MULTIPROCESSING',\n 'value': '0',\n 'value_from': None},\n {'name': 'USER',\n 'value': 'nonroot',\n 'value_from': None},\n {'name': 'TORCHINDUCTOR_CACHE_DIR',\n 'value': '/tmp/torchinductor-cache',\n 'value_from': None},\n {'name': 'HF_HUB_CACHE',\n 'value': '/models',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': {'post_start': None,\n 'pre_stop': {'_exec': {'command': ['/bin/sleep',\n '15']},\n 'http_get': None,\n 'sleep': None,\n 'tcp_socket': None}},\n 'liveness_probe': {'_exec': None,\n 'f [e2e-llm-inference-service] ailure_threshold': 10,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'name': 'main',\n 'ports': [{'container_port': 8000,\n 'host_ip': None,\n 'host_port': None,\n 'name': None,\n 'protocol': 'TCP'}],\n 'readiness_probe': {'_exec': None,\n 'failure_threshold': 2,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 1,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': True,\n 'run_as_group': None,\n 'run_as_non_root': True,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n [e2e-llm-inference-service] 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'windows_options': None},\n 'startup_probe': {'_exec': None,\n 'failure_threshold': 60,\n 'grpc': None,\n 'http_get': {'host': None,\n 'http_headers': None,\n 'path': '/health',\n 'port': 8000,\n 'scheme': 'HTTPS'},\n 'initial_delay_seconds': None,\n 'period_seconds': 10,\n 'success_threshold': 1,\n 'tcp_socket': None,\n 'termination_grace_period_seconds': None,\n 'timeout_seconds': 1},\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/home',\n 'mount_propagation': None,\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/tmp',\n 'mount_propagation': None,\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/dev/shm',\n 'mount_propagation': None,\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/models',\n 'mount_propagation': None,\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n [e2e-llm-inference-service] {'mount_path': '/var/run/kserve/tls',\n 'mount_propagation': None,\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'dns_config': None,\n 'dns_policy': 'ClusterFirst',\n 'enable_service_links': True,\n 'ephemeral_containers': None,\n 'host_aliases': None,\n 'host_ipc': None,\n 'host_network': None,\n 'host_pid': None,\n 'host_users': None,\n 'hostname': None,\n 'image_pull_secrets': [{'name': 'default-dockercfg-x9l7t'}],\n 'init_containers': [{'args': ['hf://facebook/opt-125m',\n '/mnt/models'],\n 'command': None,\n 'env': [{'name': 'AWS_ACCESS_KEY_ID',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_ACCESS_KEY_ID',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n {'name': 'AWS_SECRET_ACCESS_KEY',\n 'value': None,\n 'value_from': {'config_map_key_ref': None,\n 'field_ref': None,\n 'resource_field_ref': None,\n 'secret_key_ref': {'key': 'AWS_SECRET_ACCESS_KEY',\n 'name': 'seaweedfs-s3-creds',\n 'optional': None}}},\n [e2e-llm-inference-service] {'name': 'S3_USE_HTTPS',\n 'value': '0',\n 'value_from': None},\n {'name': 'S3_ENDPOINT',\n 'value': 's3-service.kserve:8333',\n 'value_from': None},\n {'name': 'AWS_ENDPOINT_URL',\n 'value': 'http://s3-service.kserve:8333',\n 'value_from': None},\n {'name': 'S3_VERIFY_SSL',\n 'value': '0',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE',\n 'value': '/etc/ssl/custom-certs/cabundle.crt',\n 'value_from': None},\n {'name': 'AWS_CA_BUNDLE_CONFIGMAP',\n 'value': 'odh-kserve-custom-ca-bundle',\n 'value_from': None},\n {'name': 'HF_HUB_ENABLE_HF_TRANSFER',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_HIGH_PERFORMANCE',\n 'value': '1',\n 'value_from': None},\n {'name': 'HF_XET_NUM_CONCURRENT_RANGE_GETS',\n 'value': '8',\n 'value_from': None}],\n 'env_from': None,\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_pull_policy': 'IfNotPresent',\n 'lifecycle': None,\n 'liveness_probe': None,\n 'name': 'storage-initializer',\n 'ports': None,\n 'readiness_probe': None,\n 'resize_policy': None,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_policy': None,\n 'security_context': {'allow_privilege_escalation': False,\n 'app_armor_profile': None,\n 'capabilities': {'add': None,\n 'drop': ['ALL']},\n 'privileged': None,\n 'proc_mount': None,\n 'read_only_root_filesystem': None,\n 'run_as_group': None,\n 'run_as_non_root': Tr [e2e-llm-inference-service] ue,\n 'run_as_user': 1000690000,\n 'se_linux_options': None,\n 'seccomp_profile': None,\n 'windows_options': None},\n 'startup_probe': None,\n 'stdin': None,\n 'stdin_once': None,\n 'termination_message_path': '/dev/termination-log',\n 'termination_message_policy': 'FallbackToLogsOnError',\n 'tty': None,\n 'volume_devices': None,\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'mount_propagation': None,\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'mount_propagation': None,\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': None,\n 'sub_path': None,\n 'sub_path_expr': None}],\n 'working_dir': None}],\n 'node_name': 'ip-10-0-134-1.ec2.internal',\n 'node_selector': None,\n 'os': None,\n 'overhead': None,\n 'preemption_policy': 'PreemptLowerPriority',\n 'priority': 0,\n 'priority_class_name': None,\n 'readiness_gates': None,\n 'resource_claims': None,\n 'resources': None,\n 'restart_policy': 'Always',\n 'runtime_class_name': None,\n 'scheduler_name': 'default-scheduler',\n 'scheduling_gates': None,\n 'security_context': {'app_armor_profile': None,\n 'fs_group': 1000690000,\n 'fs_group_change_policy': None,\n 'run_as_group': None,\n 'run_as_non_root': None,\n 'run_as_user': None,\n 'se_linux_change_policy': None,\n 'se_linux_options': {'level': 's0:c26,c20',\n 'role': None,\n 'type': None,\n 'user': None},\n 'seccomp_profile': {'localhost_profile': None,\n 'type': 'RuntimeDefault'},\n 'supplemental_groups': None,\n 'supplemental_groups_policy': None,\n 'sysctls': None,\n [e2e-llm-inference-service] 'windows_options': None},\n 'service_account': 'default',\n 'service_account_name': 'default',\n 'set_hostname_as_fqdn': None,\n 'share_process_namespace': None,\n 'subdomain': None,\n 'termination_grace_period_seconds': 60,\n 'tolerations': [{'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/not-ready',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoExecute',\n 'key': 'node.kubernetes.io/unreachable',\n 'operator': 'Exists',\n 'toleration_seconds': 300,\n 'value': None},\n {'effect': 'NoSchedule',\n 'key': 'node.kubernetes.io/memory-pressure',\n 'operator': 'Exists',\n 'toleration_seconds': None,\n 'value': None}],\n 'topology_spread_constraints': None,\n 'volumes': [{'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'home',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': 'Memory',\n 'size_limit': '1Gi'},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n [e2e-llm-inference-service] 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'dshm',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'model-cache',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tmp-dir',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'project [e2e-llm-inference-service] ed': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'tls-certs',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': {'default_mode': 420,\n 'items': None,\n 'optional': None,\n 'secret_name': 'tls-verification-test-kserve-self-signed-certs'},\n 'storageos': None,\n 'vsphere_volume': None},\n {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': {'medium': None,\n 'size_limit': None},\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kserve-provision-location',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': None,\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None},\n [e2e-llm-inference-service] {'aws_elastic_block_store': None,\n 'azure_disk': None,\n 'azure_file': None,\n 'cephfs': None,\n 'cinder': None,\n 'config_map': None,\n 'csi': None,\n 'downward_api': None,\n 'empty_dir': None,\n 'ephemeral': None,\n 'fc': None,\n 'flex_volume': None,\n 'flocker': None,\n 'gce_persistent_disk': None,\n 'git_repo': None,\n 'glusterfs': None,\n 'host_path': None,\n 'image': None,\n 'iscsi': None,\n 'name': 'kube-api-access-2pvc4',\n 'nfs': None,\n 'persistent_volume_claim': None,\n 'photon_persistent_disk': None,\n 'portworx_volume': None,\n 'projected': {'default_mode': 420,\n 'sources': [{'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': {'audience': None,\n 'expiration_seconds': 3607,\n 'path': 'token'}},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'ca.crt',\n 'mode': None,\n 'path': 'ca.crt'}],\n 'name': 'kube-root-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': None,\n 'downward_api': {'items': [{'field_ref': {'api_version': 'v1',\n 'field_path': 'metadata.namespace'},\n 'mode': None,\n 'path': 'namespace',\n 'resource_field_ref': None}]},\n 'secret': None,\n 'service_account_token': None},\n {'cluster_trust_bundle': None,\n 'config_map': {'items': [{'key': 'service-ca.crt',\n [e2e-llm-inference-service] 'mode': None,\n 'path': 'service-ca.crt'}],\n 'name': 'openshift-service-ca.crt',\n 'optional': None},\n 'downward_api': None,\n 'secret': None,\n 'service_account_token': None}]},\n 'quobyte': None,\n 'rbd': None,\n 'scale_io': None,\n 'secret': None,\n 'storageos': None,\n 'vsphere_volume': None}]},\n 'status': {'conditions': [{'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 27, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodReadyToStartContainers'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Initialized'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'Ready'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 13, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'ContainersReady'},\n {'last_probe_time': None,\n 'last_transition_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal()),\n 'message': None,\n 'reason': None,\n 'status': 'True',\n 'type': 'PodScheduled'}],\n 'container_statuses': [{'allocated_resources': {'cpu': '200m',\n 'memory': '2Gi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://5c346b202f8fc8d453e8879b38ce756aa0255b2f25b1496e10f2a0c4c75d6e45',\n 'image': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0',\n 'image_id': 'public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo@sha256:afb39fca138b51d019d986229d546531b45a2a3deb73bcf59bd42406e13fbba0',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n [e2e-llm-inference-service] 'name': 'main',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '2',\n 'memory': '7Gi'},\n 'requests': {'cpu': '200m',\n 'memory': '2Gi'}},\n 'restart_count': 0,\n 'started': True,\n 'state': {'running': {'started_at': datetime.datetime(2026, 7, 2, 14, 11, 31, tzinfo=tzlocal())},\n 'terminated': None,\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/home',\n 'name': 'home',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/tmp',\n 'name': 'tmp-dir',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/dev/shm',\n 'name': 'dshm',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/models',\n 'name': 'model-cache',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/kserve/tls',\n 'name': 'tls-certs',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}],\n 'ephemeral_container_statuses': None,\n 'host_i_ps': [{'ip': '10.0.134.1'}],\n [e2e-llm-inference-service] 'host_ip': '10.0.134.1',\n 'init_container_statuses': [{'allocated_resources': {'cpu': '100m',\n 'memory': '100Mi'},\n 'allocated_resources_status': None,\n 'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'image': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'image_id': 'quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b',\n 'last_state': {'running': None,\n 'terminated': None,\n 'waiting': None},\n 'name': 'storage-initializer',\n 'ready': True,\n 'resources': {'claims': None,\n 'limits': {'cpu': '1',\n 'memory': '24Gi'},\n 'requests': {'cpu': '100m',\n 'memory': '100Mi'}},\n 'restart_count': 0,\n 'started': False,\n 'state': {'running': None,\n 'terminated': {'container_id': 'cri-o://c6805dfd8caa60cae5e5e2ff7272901d954261ebc52878c60363d48bf5531aa8',\n 'exit_code': 0,\n 'finished_at': datetime.datetime(2026, 7, 2, 14, 11, 30, tzinfo=tzlocal()),\n 'message': None,\n 'reason': 'Completed',\n 'signal': None,\n 'started_at': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())},\n 'waiting': None},\n 'user': {'linux': {'gid': 0,\n 'supplemental_groups': [0,\n 1000690000],\n 'uid': 1000690000}},\n 'volume_mounts': [{'mount_path': '/mnt/models',\n 'name': 'kserve-provision-location',\n 'read_only': None,\n 'recursive_read_only': None},\n {'mount_path': '/var/run/secrets/kubernetes.io/serviceaccount',\n 'name': 'kube-api-access-2pvc4',\n 'read_only': True,\n 'recursive_read_only': 'Disabled'}]}], [e2e-llm-inference-service] \n 'message': None,\n 'nominated_node_name': None,\n 'phase': 'Running',\n 'pod_i_ps': [{'ip': '10.132.0.64'}],\n 'pod_ip': '10.132.0.64',\n 'qos_class': 'Burstable',\n 'reason': None,\n 'resize': None,\n 'resource_claim_statuses': None,\n 'start_time': datetime.datetime(2026, 7, 2, 14, 11, 26, tzinfo=tzlocal())}}],\n 'kind': 'PodList',\n 'metadata': {'_continue': None,\n 'remaining_item_count': None,\n 'resource_version': '91419',\n 'self_link': None}}.items [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [delete_llmisvc] [2026-07-02T14:28:54.581937] end - ✅ in 20.382s [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_tls_resources] [2026-07-02T14:28:54.582074] end - ❌ 1052.419s: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/tls-verification-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] _ test_llm_inference_service[router-managed-scheduler-with-configmap-ref-workload-llmd-simulator] _ [e2e-llm-inference-service] [gw0] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), chunked = False [e2e-llm-inference-service] response_conn = [e2e-llm-inference-service] preload_content = False, decode_content = False, enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] # conn.request() calls http.client.*.request, not the method in [e2e-llm-inference-service] # urllib3.request. It also calls makefile (recv) on the socket. [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn.request( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] enforce_content_length=enforce_content_length, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # We are swallowing BrokenPipeError (errno.EPIPE) since the server is [e2e-llm-inference-service] # legitimately able to close the connection after sending a valid response. [e2e-llm-inference-service] # With this behaviour, the received response is still readable. [e2e-llm-inference-service] except BrokenPipeError: [e2e-llm-inference-service] pass [e2e-llm-inference-service] except OSError as e: [e2e-llm-inference-service] # MacOS/Linux [e2e-llm-inference-service] # EPROTOTYPE and ECONNRESET are needed on macOS [e2e-llm-inference-service] # https://erickt.github.io/blog/2014/11/19/adventures-in-debugging-a-potential-osx-kernel-bug/ [e2e-llm-inference-service] # Condition changed later to emit ECONNRESET instead of only EPROTOTYPE. [e2e-llm-inference-service] if e.errno != errno.EPROTOTYPE and e.errno != errno.ECONNRESET: [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset the timeout for the recv() on the socket [e2e-llm-inference-service] read_timeout = timeout_obj.read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn.is_closed: [e2e-llm-inference-service] # In Python 3 socket.py will catch EAGAIN and return None when you [e2e-llm-inference-service] # try and read into the file pointer created by http.client, which [e2e-llm-inference-service] # instead raises a BadStatusLine exception. Instead of catching [e2e-llm-inference-service] # the exception and assuming all BadStatusLine exceptions are read [e2e-llm-inference-service] # timeouts, check for a zero timeout before making the request. [e2e-llm-inference-service] if read_timeout == 0: [e2e-llm-inference-service] raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={read_timeout})" [e2e-llm-inference-service] ) [e2e-llm-inference-service] conn.timeout = read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] # Receive the response from the server [e2e-llm-inference-service] try: [e2e-llm-inference-service] > response = conn.getresponse() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:534: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def getresponse( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] ) -> HTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get the response from the server. [e2e-llm-inference-service] [e2e-llm-inference-service] If the HTTPConnection is in the correct state, returns an instance of HTTPResponse or of whatever object is returned by the response_class variable. [e2e-llm-inference-service] [e2e-llm-inference-service] If a request has not been sent or if a previous response has not be handled, ResponseNotReady is raised. If the HTTP response indicates that the connection should be closed, then it will be closed before the response is returned. When the connection is closed, the underlying socket is closed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Raise the same error as http.client.HTTPConnection [e2e-llm-inference-service] if self._response_options is None: [e2e-llm-inference-service] raise ResponseNotReady() [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset this attribute for being used again. [e2e-llm-inference-service] resp_options = self._response_options [e2e-llm-inference-service] self._response_options = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Since the connection's timeout value may have been updated [e2e-llm-inference-service] # we need to set the timeout on the socket. [e2e-llm-inference-service] self.sock.settimeout(self.timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] # This is needed here to avoid circular import errors [e2e-llm-inference-service] from .response import HTTPResponse [e2e-llm-inference-service] [e2e-llm-inference-service] # Save a reference to the shutdown function before ownership is passed [e2e-llm-inference-service] # to httplib_response [e2e-llm-inference-service] # TODO should we implement it everywhere? [e2e-llm-inference-service] _shutdown = getattr(self.sock, "shutdown", None) [e2e-llm-inference-service] [e2e-llm-inference-service] # Get the response from http.client.HTTPConnection [e2e-llm-inference-service] > httplib_response = super().getresponse() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:571: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def getresponse(self): [e2e-llm-inference-service] """Get the response from the server. [e2e-llm-inference-service] [e2e-llm-inference-service] If the HTTPConnection is in the correct state, returns an [e2e-llm-inference-service] instance of HTTPResponse or of whatever object is returned by [e2e-llm-inference-service] the response_class variable. [e2e-llm-inference-service] [e2e-llm-inference-service] If a request has not been sent or if a previous response has [e2e-llm-inference-service] not be handled, ResponseNotReady is raised. If the HTTP [e2e-llm-inference-service] response indicates that the connection should be closed, then [e2e-llm-inference-service] it will be closed before the response is returned. When the [e2e-llm-inference-service] connection is closed, the underlying socket is closed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] # if a prior response has been completed, then forget about it. [e2e-llm-inference-service] if self.__response and self.__response.isclosed(): [e2e-llm-inference-service] self.__response = None [e2e-llm-inference-service] [e2e-llm-inference-service] # if a prior response exists, then it must be completed (otherwise, we [e2e-llm-inference-service] # cannot read this response's header to determine the connection-close [e2e-llm-inference-service] # behavior) [e2e-llm-inference-service] # [e2e-llm-inference-service] # note: if a prior response existed, but was connection-close, then the [e2e-llm-inference-service] # socket and response were made independent of this HTTPConnection [e2e-llm-inference-service] # object since a new request requires that we open a whole new [e2e-llm-inference-service] # connection [e2e-llm-inference-service] # [e2e-llm-inference-service] # this means the prior response had one of two states: [e2e-llm-inference-service] # 1) will_close: this connection was reset and the prior socket and [e2e-llm-inference-service] # response operate independently [e2e-llm-inference-service] # 2) persistent: the response was retained and we await its [e2e-llm-inference-service] # isclosed() status to become true. [e2e-llm-inference-service] # [e2e-llm-inference-service] if self.__state != _CS_REQ_SENT or self.__response: [e2e-llm-inference-service] raise ResponseNotReady(self.__state) [e2e-llm-inference-service] [e2e-llm-inference-service] if self.debuglevel > 0: [e2e-llm-inference-service] response = self.response_class(self.sock, self.debuglevel, [e2e-llm-inference-service] method=self._method) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = self.response_class(self.sock, method=self._method) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > response.begin() [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:1395: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def begin(self): [e2e-llm-inference-service] if self.headers is not None: [e2e-llm-inference-service] # we've already started reading the response [e2e-llm-inference-service] return [e2e-llm-inference-service] [e2e-llm-inference-service] # read until we get a non-100 response [e2e-llm-inference-service] while True: [e2e-llm-inference-service] > version, status, reason = self._read_status() [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:325: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def _read_status(self): [e2e-llm-inference-service] > line = str(self.fp.readline(_MAXLINE + 1), "iso-8859-1") [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/http/client.py:286: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] b = [e2e-llm-inference-service] [e2e-llm-inference-service] def readinto(self, b): [e2e-llm-inference-service] """Read up to len(b) bytes into the writable buffer *b* and return [e2e-llm-inference-service] the number of bytes read. If the socket is non-blocking and no bytes [e2e-llm-inference-service] are available, None is returned. [e2e-llm-inference-service] [e2e-llm-inference-service] If *b* is non-empty, a 0 return value indicates that the connection [e2e-llm-inference-service] was shutdown at the other end. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self._checkClosed() [e2e-llm-inference-service] self._checkReadable() [e2e-llm-inference-service] if self._timeout_occurred: [e2e-llm-inference-service] raise OSError("cannot read from timed out object") [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return self._sock.recv_into(b) [e2e-llm-inference-service] E TimeoutError: timed out [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/socket.py:718: TimeoutError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] > response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:787: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), chunked = False [e2e-llm-inference-service] response_conn = [e2e-llm-inference-service] preload_content = False, decode_content = False, enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] # conn.request() calls http.client.*.request, not the method in [e2e-llm-inference-service] # urllib3.request. It also calls makefile (recv) on the socket. [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn.request( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] enforce_content_length=enforce_content_length, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # We are swallowing BrokenPipeError (errno.EPIPE) since the server is [e2e-llm-inference-service] # legitimately able to close the connection after sending a valid response. [e2e-llm-inference-service] # With this behaviour, the received response is still readable. [e2e-llm-inference-service] except BrokenPipeError: [e2e-llm-inference-service] pass [e2e-llm-inference-service] except OSError as e: [e2e-llm-inference-service] # MacOS/Linux [e2e-llm-inference-service] # EPROTOTYPE and ECONNRESET are needed on macOS [e2e-llm-inference-service] # https://erickt.github.io/blog/2014/11/19/adventures-in-debugging-a-potential-osx-kernel-bug/ [e2e-llm-inference-service] # Condition changed later to emit ECONNRESET instead of only EPROTOTYPE. [e2e-llm-inference-service] if e.errno != errno.EPROTOTYPE and e.errno != errno.ECONNRESET: [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # Reset the timeout for the recv() on the socket [e2e-llm-inference-service] read_timeout = timeout_obj.read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn.is_closed: [e2e-llm-inference-service] # In Python 3 socket.py will catch EAGAIN and return None when you [e2e-llm-inference-service] # try and read into the file pointer created by http.client, which [e2e-llm-inference-service] # instead raises a BadStatusLine exception. Instead of catching [e2e-llm-inference-service] # the exception and assuming all BadStatusLine exceptions are read [e2e-llm-inference-service] # timeouts, check for a zero timeout before making the request. [e2e-llm-inference-service] if read_timeout == 0: [e2e-llm-inference-service] raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={read_timeout})" [e2e-llm-inference-service] ) [e2e-llm-inference-service] conn.timeout = read_timeout [e2e-llm-inference-service] [e2e-llm-inference-service] # Receive the response from the server [e2e-llm-inference-service] try: [e2e-llm-inference-service] response = conn.getresponse() [e2e-llm-inference-service] except (BaseSSLError, OSError) as e: [e2e-llm-inference-service] > self._raise_timeout(err=e, url=url, timeout_value=read_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:536: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] err = TimeoutError('timed out') [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] timeout_value = 60 [e2e-llm-inference-service] [e2e-llm-inference-service] def _raise_timeout( [e2e-llm-inference-service] self, [e2e-llm-inference-service] err: BaseSSLError | OSError | SocketTimeout, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] timeout_value: _TYPE_TIMEOUT | None, [e2e-llm-inference-service] ) -> None: [e2e-llm-inference-service] """Is the error actually a timeout? Will raise a ReadTimeout or pass""" [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(err, SocketTimeout): [e2e-llm-inference-service] > raise ReadTimeoutError( [e2e-llm-inference-service] self, url, f"Read timed out. (read timeout={timeout_value})" [e2e-llm-inference-service] ) from err [e2e-llm-inference-service] E urllib3.exceptions.ReadTimeoutError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:367: ReadTimeoutError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = , stream = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), verify = '/tmp/ca.crt' [e2e-llm-inference-service] cert = None, proxies = OrderedDict() [e2e-llm-inference-service] [e2e-llm-inference-service] def send( [e2e-llm-inference-service] self, request, stream=False, timeout=None, verify=True, cert=None, proxies=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Sends PreparedRequest object. Returns Response object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param request: The :class:`PreparedRequest ` being sent. [e2e-llm-inference-service] :param stream: (optional) Whether to stream the request content. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple or urllib3 Timeout object [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether [e2e-llm-inference-service] we verify the server's TLS certificate, or a string, in which case it [e2e-llm-inference-service] must be a path to a CA bundle to use [e2e-llm-inference-service] :param cert: (optional) Any user-provided SSL certificate to be trusted. [e2e-llm-inference-service] :param proxies: (optional) The proxies dictionary to apply to the request. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn = self.get_connection_with_tls_context( [e2e-llm-inference-service] request, verify, proxies=proxies, cert=cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] except LocationValueError as e: [e2e-llm-inference-service] raise InvalidURL(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] self.cert_verify(conn, request.url, verify, cert) [e2e-llm-inference-service] url = self.request_url(request, proxies) [e2e-llm-inference-service] self.add_headers( [e2e-llm-inference-service] request, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] verify=verify, [e2e-llm-inference-service] cert=cert, [e2e-llm-inference-service] proxies=proxies, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] chunked = not (request.body is None or "Content-Length" in request.headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(timeout, tuple): [e2e-llm-inference-service] try: [e2e-llm-inference-service] connect, read = timeout [e2e-llm-inference-service] timeout = TimeoutSauce(connect=connect, read=read) [e2e-llm-inference-service] except ValueError: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, " [e2e-llm-inference-service] f"or a single float to set both timeouts to the same value." [e2e-llm-inference-service] ) [e2e-llm-inference-service] elif isinstance(timeout, TimeoutSauce): [e2e-llm-inference-service] pass [e2e-llm-inference-service] else: [e2e-llm-inference-service] timeout = TimeoutSauce(connect=timeout, read=timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > resp = conn.urlopen( [e2e-llm-inference-service] method=request.method, [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] body=request.body, [e2e-llm-inference-service] headers=request.headers, [e2e-llm-inference-service] redirect=False, [e2e-llm-inference-service] assert_same_host=False, [e2e-llm-inference-service] preload_content=False, [e2e-llm-inference-service] decode_content=False, [e2e-llm-inference-service] retries=self.max_retries, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/adapters.py:667: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=7, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=6, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=5, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=4, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=3, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=2, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=1, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = RemoteDisconnected('Remote end closed connection without response') [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] body = b'{"model": "facebook/opt-125m", "prompt": "KServe is a", "max_tokens": 20}' [e2e-llm-inference-service] headers = {'User-Agent': 'python-requests/2.32.3', 'Accept-Encoding': 'gzip, deflate', 'Accept': '*/*', 'Connection': 'keep-alive', 'Content-Type': 'application/json', 'Content-Length': '73'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), pool_timeout = None [e2e-llm-inference-service] release_conn = False, chunked = False, body_pos = None, preload_content = False [e2e-llm-inference-service] decode_content = False, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions', query=None, fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] > retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:841: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] method = 'POST' [e2e-llm-inference-service] url = '/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] response = None [e2e-llm-inference-service] error = ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)") [e2e-llm-inference-service] _pool = [e2e-llm-inference-service] _stacktrace = [e2e-llm-inference-service] [e2e-llm-inference-service] def increment( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str | None = None, [e2e-llm-inference-service] url: str | None = None, [e2e-llm-inference-service] response: BaseHTTPResponse | None = None, [e2e-llm-inference-service] error: Exception | None = None, [e2e-llm-inference-service] _pool: ConnectionPool | None = None, [e2e-llm-inference-service] _stacktrace: TracebackType | None = None, [e2e-llm-inference-service] ) -> Self: [e2e-llm-inference-service] """Return a new Retry object with incremented retry counters. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response: A response object, or None, if the server did not [e2e-llm-inference-service] return a response. [e2e-llm-inference-service] :type response: :class:`~urllib3.response.BaseHTTPResponse` [e2e-llm-inference-service] :param Exception error: An error encountered during the request, or [e2e-llm-inference-service] None if the response was received successfully. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: A new ``Retry`` object. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if self.total is False and error: [e2e-llm-inference-service] # Disabled, indicate to re-raise the error. [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] [e2e-llm-inference-service] total = self.total [e2e-llm-inference-service] if total is not None: [e2e-llm-inference-service] total -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] connect = self.connect [e2e-llm-inference-service] read = self.read [e2e-llm-inference-service] redirect = self.redirect [e2e-llm-inference-service] status_count = self.status [e2e-llm-inference-service] other = self.other [e2e-llm-inference-service] cause = "unknown" [e2e-llm-inference-service] status = None [e2e-llm-inference-service] redirect_location = None [e2e-llm-inference-service] [e2e-llm-inference-service] if error and self._is_connection_error(error): [e2e-llm-inference-service] # Connect retry? [e2e-llm-inference-service] if connect is False: [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif connect is not None: [e2e-llm-inference-service] connect -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error and self._is_read_error(error): [e2e-llm-inference-service] # Read retry? [e2e-llm-inference-service] if read is False or method is None or not self._is_method_retryable(method): [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif read is not None: [e2e-llm-inference-service] read -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error: [e2e-llm-inference-service] # Other retry? [e2e-llm-inference-service] if other is not None: [e2e-llm-inference-service] other -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif response and response.get_redirect_location(): [e2e-llm-inference-service] # Redirect retry? [e2e-llm-inference-service] if redirect is not None: [e2e-llm-inference-service] redirect -= 1 [e2e-llm-inference-service] cause = "too many redirects" [e2e-llm-inference-service] response_redirect_location = response.get_redirect_location() [e2e-llm-inference-service] if response_redirect_location: [e2e-llm-inference-service] redirect_location = response_redirect_location [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Incrementing because of a server error like a 500 in [e2e-llm-inference-service] # status_forcelist and the given method is in the allowed_methods [e2e-llm-inference-service] cause = ResponseError.GENERIC_ERROR [e2e-llm-inference-service] if response and response.status: [e2e-llm-inference-service] if status_count is not None: [e2e-llm-inference-service] status_count -= 1 [e2e-llm-inference-service] cause = ResponseError.SPECIFIC_ERROR.format(status_code=response.status) [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] history = self.history + ( [e2e-llm-inference-service] RequestHistory(method, url, error, status, redirect_location), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] new_retry = self.new( [e2e-llm-inference-service] total=total, [e2e-llm-inference-service] connect=connect, [e2e-llm-inference-service] read=read, [e2e-llm-inference-service] redirect=redirect, [e2e-llm-inference-service] status=status_count, [e2e-llm-inference-service] other=other, [e2e-llm-inference-service] history=history, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if new_retry.is_exhausted(): [e2e-llm-inference-service] reason = error or ResponseError(cause) [e2e-llm-inference-service] > raise MaxRetryError(_pool, url, reason) from reason # type: ignore[arg-type] [e2e-llm-inference-service] E urllib3.exceptions.MaxRetryError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/util/retry.py:519: MaxRetryError [e2e-llm-inference-service] [e2e-llm-inference-service] During handling of the above exception, another exception occurred: [e2e-llm-inference-service] [e2e-llm-inference-service] def get_successful_response(): [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_case.url_getter: [e2e-llm-inference-service] service_url = test_case.url_getter(kserve_client, test_case.llm_service) [e2e-llm-inference-service] else: [e2e-llm-inference-service] service_url = get_llm_service_url(kserve_client, test_case.llm_service) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to get service URL: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] model_url = service_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] headers = {"Content-Type": "application/json"} [e2e-llm-inference-service] if extra_headers: [e2e-llm-inference-service] headers.update(extra_headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if test_case.payload_formatter is not None: [e2e-llm-inference-service] test_payload = test_case.payload_formatter(test_case) [e2e-llm-inference-service] elif test_case.prompt is not None: [e2e-llm-inference-service] test_payload = { [e2e-llm-inference-service] "model": test_case.model_name [e2e-llm-inference-service] if not extra_headers or MODEL_ROUTING_HEADER not in extra_headers [e2e-llm-inference-service] else extra_headers[MODEL_ROUTING_HEADER], [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] else: [e2e-llm-inference-service] test_payload = None [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Calling LLM service at {model_url} with payload {test_payload}") [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_payload is not None: [e2e-llm-inference-service] > response = post_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] json_data=test_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1095: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] [e2e-llm-inference-service] def post_with_retry( [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] *, [e2e-llm-inference-service] headers: Dict = None, [e2e-llm-inference-service] json_data: Union[Dict, List] = None, [e2e-llm-inference-service] data: Union[str, bytes] = None, [e2e-llm-inference-service] stream: bool = False, [e2e-llm-inference-service] timeout: float = None, [e2e-llm-inference-service] total_retries: int = DEFAULT_RETRY_TOTAL, [e2e-llm-inference-service] backoff_factor: float = DEFAULT_RETRY_BACKOFF_FACTOR, [e2e-llm-inference-service] retry_status_codes=DEFAULT_RETRY_STATUS_CODES, [e2e-llm-inference-service] ) -> requests.Response: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Send POST request with retries for transient HTTP and network failures. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if json_data is not None and data is not None: [e2e-llm-inference-service] raise ValueError("Only one of json_data or data can be provided.") [e2e-llm-inference-service] [e2e-llm-inference-service] with _retry_session( [e2e-llm-inference-service] ["POST"], total_retries, backoff_factor, retry_status_codes [e2e-llm-inference-service] ) as session: [e2e-llm-inference-service] > return session.post( [e2e-llm-inference-service] url, [e2e-llm-inference-service] json=json_data, [e2e-llm-inference-service] data=data, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] common/http_retry.py:70: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] data = None [e2e-llm-inference-service] json = {'max_tokens': 20, 'model': 'facebook/opt-125m', 'prompt': 'KServe is a'} [e2e-llm-inference-service] kwargs = {'headers': {'Content-Type': 'application/json'}, 'stream': False, 'timeout': 60} [e2e-llm-inference-service] [e2e-llm-inference-service] def post(self, url, data=None, json=None, **kwargs): [e2e-llm-inference-service] r"""Sends a POST request. Returns :class:`Response` object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: URL for the new :class:`Request` object. [e2e-llm-inference-service] :param data: (optional) Dictionary, list of tuples, bytes, or file-like [e2e-llm-inference-service] object to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param json: (optional) json to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param \*\*kwargs: Optional arguments that ``request`` takes. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.request("POST", url, data=data, json=json, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:637: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = , method = 'POST' [e2e-llm-inference-service] url = 'http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions' [e2e-llm-inference-service] params = None, data = None, headers = {'Content-Type': 'application/json'} [e2e-llm-inference-service] cookies = None, files = None, auth = None, timeout = 60, allow_redirects = True [e2e-llm-inference-service] proxies = {}, hooks = None, stream = False, verify = None, cert = None [e2e-llm-inference-service] json = {'max_tokens': 20, 'model': 'facebook/opt-125m', 'prompt': 'KServe is a'} [e2e-llm-inference-service] [e2e-llm-inference-service] def request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] params=None, [e2e-llm-inference-service] data=None, [e2e-llm-inference-service] headers=None, [e2e-llm-inference-service] cookies=None, [e2e-llm-inference-service] files=None, [e2e-llm-inference-service] auth=None, [e2e-llm-inference-service] timeout=None, [e2e-llm-inference-service] allow_redirects=True, [e2e-llm-inference-service] proxies=None, [e2e-llm-inference-service] hooks=None, [e2e-llm-inference-service] stream=None, [e2e-llm-inference-service] verify=None, [e2e-llm-inference-service] cert=None, [e2e-llm-inference-service] json=None, [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Constructs a :class:`Request `, prepares it and sends it. [e2e-llm-inference-service] Returns :class:`Response ` object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: method for the new :class:`Request` object. [e2e-llm-inference-service] :param url: URL for the new :class:`Request` object. [e2e-llm-inference-service] :param params: (optional) Dictionary or bytes to be sent in the query [e2e-llm-inference-service] string for the :class:`Request`. [e2e-llm-inference-service] :param data: (optional) Dictionary, list of tuples, bytes, or file-like [e2e-llm-inference-service] object to send in the body of the :class:`Request`. [e2e-llm-inference-service] :param json: (optional) json to send in the body of the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param headers: (optional) Dictionary of HTTP Headers to send with the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param cookies: (optional) Dict or CookieJar object to send with the [e2e-llm-inference-service] :class:`Request`. [e2e-llm-inference-service] :param files: (optional) Dictionary of ``'filename': file-like-objects`` [e2e-llm-inference-service] for multipart encoding upload. [e2e-llm-inference-service] :param auth: (optional) Auth tuple or callable to enable [e2e-llm-inference-service] Basic/Digest/Custom HTTP Auth. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple [e2e-llm-inference-service] :param allow_redirects: (optional) Set to True by default. [e2e-llm-inference-service] :type allow_redirects: bool [e2e-llm-inference-service] :param proxies: (optional) Dictionary mapping protocol or protocol and [e2e-llm-inference-service] hostname to the URL of the proxy. [e2e-llm-inference-service] :param hooks: (optional) Dictionary mapping hook name to one event or [e2e-llm-inference-service] list of events, event must be callable. [e2e-llm-inference-service] :param stream: (optional) whether to immediately download the response [e2e-llm-inference-service] content. Defaults to ``False``. [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether we verify [e2e-llm-inference-service] the server's TLS certificate, or a string, in which case it must be a path [e2e-llm-inference-service] to a CA bundle to use. Defaults to ``True``. When set to [e2e-llm-inference-service] ``False``, requests will accept any TLS certificate presented by [e2e-llm-inference-service] the server, and will ignore hostname mismatches and/or expired [e2e-llm-inference-service] certificates, which will make your application vulnerable to [e2e-llm-inference-service] man-in-the-middle (MitM) attacks. Setting verify to ``False`` [e2e-llm-inference-service] may be useful during local development or testing. [e2e-llm-inference-service] :param cert: (optional) if String, path to ssl client cert file (.pem). [e2e-llm-inference-service] If Tuple, ('cert', 'key') pair. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Create the Request. [e2e-llm-inference-service] req = Request( [e2e-llm-inference-service] method=method.upper(), [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] files=files, [e2e-llm-inference-service] data=data or {}, [e2e-llm-inference-service] json=json, [e2e-llm-inference-service] params=params or {}, [e2e-llm-inference-service] auth=auth, [e2e-llm-inference-service] cookies=cookies, [e2e-llm-inference-service] hooks=hooks, [e2e-llm-inference-service] ) [e2e-llm-inference-service] prep = self.prepare_request(req) [e2e-llm-inference-service] [e2e-llm-inference-service] proxies = proxies or {} [e2e-llm-inference-service] [e2e-llm-inference-service] settings = self.merge_environment_settings( [e2e-llm-inference-service] prep.url, proxies, stream, verify, cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Send the request. [e2e-llm-inference-service] send_kwargs = { [e2e-llm-inference-service] "timeout": timeout, [e2e-llm-inference-service] "allow_redirects": allow_redirects, [e2e-llm-inference-service] } [e2e-llm-inference-service] send_kwargs.update(settings) [e2e-llm-inference-service] > resp = self.send(prep, **send_kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:589: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = [e2e-llm-inference-service] kwargs = {'cert': None, 'proxies': OrderedDict(), 'stream': False, 'timeout': 60, ...} [e2e-llm-inference-service] allow_redirects = True, stream = False, hooks = {'response': []} [e2e-llm-inference-service] adapter = [e2e-llm-inference-service] start = 1783001668.6000636 [e2e-llm-inference-service] [e2e-llm-inference-service] def send(self, request, **kwargs): [e2e-llm-inference-service] """Send a given PreparedRequest. [e2e-llm-inference-service] [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] # Set defaults that the hooks can utilize to ensure they always have [e2e-llm-inference-service] # the correct parameters to reproduce the previous request. [e2e-llm-inference-service] kwargs.setdefault("stream", self.stream) [e2e-llm-inference-service] kwargs.setdefault("verify", self.verify) [e2e-llm-inference-service] kwargs.setdefault("cert", self.cert) [e2e-llm-inference-service] if "proxies" not in kwargs: [e2e-llm-inference-service] kwargs["proxies"] = resolve_proxies(request, self.proxies, self.trust_env) [e2e-llm-inference-service] [e2e-llm-inference-service] # It's possible that users might accidentally send a Request object. [e2e-llm-inference-service] # Guard against that specific failure case. [e2e-llm-inference-service] if isinstance(request, Request): [e2e-llm-inference-service] raise ValueError("You can only send PreparedRequests.") [e2e-llm-inference-service] [e2e-llm-inference-service] # Set up variables needed for resolve_redirects and dispatching of hooks [e2e-llm-inference-service] allow_redirects = kwargs.pop("allow_redirects", True) [e2e-llm-inference-service] stream = kwargs.get("stream") [e2e-llm-inference-service] hooks = request.hooks [e2e-llm-inference-service] [e2e-llm-inference-service] # Get the appropriate adapter to use [e2e-llm-inference-service] adapter = self.get_adapter(url=request.url) [e2e-llm-inference-service] [e2e-llm-inference-service] # Start time (approximately) of the request [e2e-llm-inference-service] start = preferred_clock() [e2e-llm-inference-service] [e2e-llm-inference-service] # Send the request [e2e-llm-inference-service] > r = adapter.send(request, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/sessions.py:703: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] request = , stream = False [e2e-llm-inference-service] timeout = Timeout(connect=60, read=60, total=None), verify = '/tmp/ca.crt' [e2e-llm-inference-service] cert = None, proxies = OrderedDict() [e2e-llm-inference-service] [e2e-llm-inference-service] def send( [e2e-llm-inference-service] self, request, stream=False, timeout=None, verify=True, cert=None, proxies=None [e2e-llm-inference-service] ): [e2e-llm-inference-service] """Sends PreparedRequest object. Returns Response object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param request: The :class:`PreparedRequest ` being sent. [e2e-llm-inference-service] :param stream: (optional) Whether to stream the request content. [e2e-llm-inference-service] :param timeout: (optional) How long to wait for the server to send [e2e-llm-inference-service] data before giving up, as a float, or a :ref:`(connect timeout, [e2e-llm-inference-service] read timeout) ` tuple. [e2e-llm-inference-service] :type timeout: float or tuple or urllib3 Timeout object [e2e-llm-inference-service] :param verify: (optional) Either a boolean, in which case it controls whether [e2e-llm-inference-service] we verify the server's TLS certificate, or a string, in which case it [e2e-llm-inference-service] must be a path to a CA bundle to use [e2e-llm-inference-service] :param cert: (optional) Any user-provided SSL certificate to be trusted. [e2e-llm-inference-service] :param proxies: (optional) The proxies dictionary to apply to the request. [e2e-llm-inference-service] :rtype: requests.Response [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] conn = self.get_connection_with_tls_context( [e2e-llm-inference-service] request, verify, proxies=proxies, cert=cert [e2e-llm-inference-service] ) [e2e-llm-inference-service] except LocationValueError as e: [e2e-llm-inference-service] raise InvalidURL(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] self.cert_verify(conn, request.url, verify, cert) [e2e-llm-inference-service] url = self.request_url(request, proxies) [e2e-llm-inference-service] self.add_headers( [e2e-llm-inference-service] request, [e2e-llm-inference-service] stream=stream, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] verify=verify, [e2e-llm-inference-service] cert=cert, [e2e-llm-inference-service] proxies=proxies, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] chunked = not (request.body is None or "Content-Length" in request.headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(timeout, tuple): [e2e-llm-inference-service] try: [e2e-llm-inference-service] connect, read = timeout [e2e-llm-inference-service] timeout = TimeoutSauce(connect=connect, read=read) [e2e-llm-inference-service] except ValueError: [e2e-llm-inference-service] raise ValueError( [e2e-llm-inference-service] f"Invalid timeout {timeout}. Pass a (connect, read) timeout tuple, " [e2e-llm-inference-service] f"or a single float to set both timeouts to the same value." [e2e-llm-inference-service] ) [e2e-llm-inference-service] elif isinstance(timeout, TimeoutSauce): [e2e-llm-inference-service] pass [e2e-llm-inference-service] else: [e2e-llm-inference-service] timeout = TimeoutSauce(connect=timeout, read=timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] resp = conn.urlopen( [e2e-llm-inference-service] method=request.method, [e2e-llm-inference-service] url=url, [e2e-llm-inference-service] body=request.body, [e2e-llm-inference-service] headers=request.headers, [e2e-llm-inference-service] redirect=False, [e2e-llm-inference-service] assert_same_host=False, [e2e-llm-inference-service] preload_content=False, [e2e-llm-inference-service] decode_content=False, [e2e-llm-inference-service] retries=self.max_retries, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] except (ProtocolError, OSError) as err: [e2e-llm-inference-service] raise ConnectionError(err, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] except MaxRetryError as e: [e2e-llm-inference-service] if isinstance(e.reason, ConnectTimeoutError): [e2e-llm-inference-service] # TODO: Remove this in 3.0.0: see #2811 [e2e-llm-inference-service] if not isinstance(e.reason, NewConnectionError): [e2e-llm-inference-service] raise ConnectTimeout(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, ResponseError): [e2e-llm-inference-service] raise RetryError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, _ProxyError): [e2e-llm-inference-service] raise ProxyError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] if isinstance(e.reason, _SSLError): [e2e-llm-inference-service] # This branch is for urllib3 v1.22 and later. [e2e-llm-inference-service] raise SSLError(e, request=request) [e2e-llm-inference-service] [e2e-llm-inference-service] > raise ConnectionError(e, request=request) [e2e-llm-inference-service] E requests.exceptions.ConnectionError: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/requests/adapters.py:700: ConnectionError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'scheduler-with-configmap-ref', 'workload-llmd-simulator'], prompt='KServe is a'... {'name': 'workload-llmd-simulator-schedul-b2f159da'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-gateway-ref", [e2e-llm-inference-service] "router-with-managed-route", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[0], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[0]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="custom-route-timeout-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="router-with-refs-test", [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[0], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[0]], [e2e-llm-inference-service] routes=[ROUTER_ROUTES[0], ROUTER_ROUTES[1]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=["router-managed", "workload-pd-cpu", "model-fb-opt-125m"], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="custom-route-timeout-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="router-with-refs-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[1], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[1]], [e2e-llm-inference-service] routes=[ROUTER_ROUTES[2], ROUTER_ROUTES[3]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-dp-ep-gpu", [e2e-llm-inference-service] "workload-dp-ep-prefill-gpu", [e2e-llm-inference-service] "model-deepseek-v2-lite", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="Delve into the multifaceted implications of a fully disaggregated cloud architecture, specifically " [e2e-llm-inference-service] "where the compute plane (P) and the data plane (D) are independently deployed and managed for a " [e2e-llm-inference-service] "geographically distributed, high-throughput, low-latency microservices ecosystem. Beyond the " [e2e-llm-inference-service] "fundamental challenges of network latency and data consistency, elaborate on the advanced " [e2e-llm-inference-service] "considerations and trade-offs inherent in such a setup: 1. Network Architecture and Protocols: " [e2e-llm-inference-service] "How would the network fabric and underlying protocols (e.g., RDMA, custom transport layers) need to " [e2e-llm-inference-service] "evolve to support optimal performance and minimize inter-plane communication overhead, especially for " [e2e-llm-inference-service] "synchronous operations? Discuss the role of network programmability (e.g., SDN, P4) in dynamically " [e2e-llm-inference-service] "optimizing routing and traffic flow between P and D. 2. Advanced Data Consistency and Durability: " [e2e-llm-inference-service] "Explore sophisticated data consistency models (e.g., causal consistency, strong eventual consistency) " [e2e-llm-inference-service] "and their applicability in balancing performance and data integrity across a globally distributed data plane. " [e2e-llm-inference-service] "Detail strategies for ensuring data durability and fault tolerance, including multi-region replication, " [e2e-llm-inference-service] "intelligent partitioning, and recovery mechanisms in the event of partial or full plane failures. " [e2e-llm-inference-service] "3. Dynamic Resource Orchestration and Cost Optimization: Analyze how an orchestration layer would intelligently " [e2e-llm-inference-service] "manage the independent scaling of compute (P) and data (D) resources, considering fluctuating workloads, " [e2e-llm-inference-service] "cost efficiency, and performance targets (e.g., using predictive analytics for resource provisioning). " [e2e-llm-inference-service] "Discuss mechanisms for dynamically reallocating compute nodes to different data partitions based on " [e2e-llm-inference-service] "workload patterns and data locality, potentially involving live migration strategies. " [e2e-llm-inference-service] "4. Security and Compliance in a Distributed Landscape: Address the enhanced security perimeter " [e2e-llm-inference-service] "challenges, including securing communication channels between P and D (encryption in transit, mutual TLS), " [e2e-llm-inference-service] "fine-grained access control to data at rest and in motion, and identity management across disaggregated " [e2e-llm-inference-service] "components. Discuss how such an architecture impacts compliance with regulatory frameworks (e.g., GDPR, HIPAA) " [e2e-llm-inference-service] "concerning data sovereignty, privacy, and auditability. 5. Operational Complexity and Observability: " [e2e-llm-inference-service] "Examine the increased complexity in monitoring, logging, and tracing across highly decoupled compute and " [e2e-llm-inference-service] "data planes. What specialized tooling and practices (e.g., distributed tracing with OpenTelemetry, advanced AIOps) " [e2e-llm-inference-service] "would be essential? How would incident response and troubleshooting differ in this disaggregated environment " [e2e-llm-inference-service] "compared to traditional integrated systems? Consider the challenges of pinpointing root causes across " [e2e-llm-inference-service] "independent failures. 6. Real-world Applicability and Future Trends: Identify specific industries " [e2e-llm-inference-service] "or use cases (e.g., high-frequency trading, IoT edge processing, large language model inference) " [e2e-llm-inference-service] "where the benefits of P/D disaggregation would strongly outweigh its complexities. " [e2e-llm-inference-service] "Conclude by speculating on emerging technologies or paradigms (e.g., serverless compute functions " [e2e-llm-inference-service] "directly interacting with object storage, in-memory disaggregation) that could further drive or " [e2e-llm-inference-service] "transform P/D disaggregation in cloud computing.", [e2e-llm-inference-service] max_tokens=2000, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_gpu, [e2e-llm-inference-service] pytest.mark.cluster_nvidia, [e2e-llm-inference-service] pytest.mark.cluster_nvidia_roce, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-no-scheduler", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.no_scheduler, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="This test simulates DP+EP that can run on CPU, the idea is to test the LWS-based deployment, " [e2e-llm-inference-service] "but without the resources requirements for DP+EP (GPUs and ROCe/IB).", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_multi_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler config tests [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-inline-config-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Chat completions endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] model_name="Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-configmap-ref", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-configmap-ref-test", [e2e-llm-inference-service] before_test=[create_scheduler_configmap], [e2e-llm-inference-service] after_test=[delete_scheduler_configmap], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-replicas", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-ha-replicas-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-custom-template", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-custom-template-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler v0.6 → v0.7 migration tests. [e2e-llm-inference-service] # Deploy v0.6-style configs and verify the controller migrates them [e2e-llm-inference-service] # so the v0.7 scheduler boots successfully. [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-pd-config-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-pd-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-nonzero-threshold-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-threshold-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Precise prefix KV cache routing test [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-precise-prefix-cache-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator-kvcache", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="precise-prefix-cache-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Models endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="data"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/chat/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — LoRA adapter [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] model_name=f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/models (base + LoRA) [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=assert_models_contains( [e2e-llm-inference-service] "facebook/opt-125m", [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] "lora-adapter-1", [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # PVC storage tests -- validate direct PVC volume mount with real vLLM serving [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_multi_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_inference_service(test_case: TestCase): # noqa: F811 [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] if not test_case.llm_service.metadata.annotations: [e2e-llm-inference-service] test_case.llm_service.metadata.annotations = {} [e2e-llm-inference-service] [e2e-llm-inference-service] test_case.llm_service.metadata.annotations[ [e2e-llm-inference-service] "security.opendatahub.io/enable-auth" [e2e-llm-inference-service] ] = "false" [e2e-llm-inference-service] prefix = test_case.log_prefix [e2e-llm-inference-service] [e2e-llm-inference-service] test_failed = False [e2e-llm-inference-service] try: [e2e-llm-inference-service] print(f"{prefix} Creating LLMInferenceService {service_name}") [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] print(f"{prefix} Waiting for LLMInferenceService {service_name} to be ready") [e2e-llm-inference-service] wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] print(f"{prefix} Waiting for model response from {service_name}") [e2e-llm-inference-service] > wait_for_model_response( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] test_case, [e2e-llm-inference-service] test_case.wait_timeout, [e2e-llm-inference-service] extra_headers=test_case.extra_headers, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:816: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, TestCase(base_refs=['router-managed', 'scheduler-wi... {'name': 'workload-llmd-simulator-schedul-b2f159da'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m'), 900) [e2e-llm-inference-service] kwargs = {'extra_headers': None}, func_name = 'wait_for_model_response' [e2e-llm-inference-service] timestamp_start = '2026-07-02T14:14:28.589094', start_time = 1783001668.5893579 [e2e-llm-inference-service] duration = 904.54163813591, timestamp_end = '2026-07-02T14:29:33.130999' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'scheduler-with-configmap-ref', 'workload-llmd-simulator'], prompt='KServe is a'... {'name': 'workload-llmd-simulator-schedul-b2f159da'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] timeout_seconds = 900, extra_headers = None [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_model_response( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] test_case: TestCase, # noqa: F811 [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] extra_headers: Optional[Dict[str, str]] = None, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def get_successful_response(): [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_case.url_getter: [e2e-llm-inference-service] service_url = test_case.url_getter(kserve_client, test_case.llm_service) [e2e-llm-inference-service] else: [e2e-llm-inference-service] service_url = get_llm_service_url(kserve_client, test_case.llm_service) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to get service URL: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] model_url = service_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] headers = {"Content-Type": "application/json"} [e2e-llm-inference-service] if extra_headers: [e2e-llm-inference-service] headers.update(extra_headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if test_case.payload_formatter is not None: [e2e-llm-inference-service] test_payload = test_case.payload_formatter(test_case) [e2e-llm-inference-service] elif test_case.prompt is not None: [e2e-llm-inference-service] test_payload = { [e2e-llm-inference-service] "model": test_case.model_name [e2e-llm-inference-service] if not extra_headers or MODEL_ROUTING_HEADER not in extra_headers [e2e-llm-inference-service] else extra_headers[MODEL_ROUTING_HEADER], [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] else: [e2e-llm-inference-service] test_payload = None [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Calling LLM service at {model_url} with payload {test_payload}") [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_payload is not None: [e2e-llm-inference-service] response = post_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] json_data=test_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = get_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] logger.error(f"❌ Failed to call model: {e}") [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to call model: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Model response is {response.status_code}: {response.text[:500]}") [e2e-llm-inference-service] [e2e-llm-inference-service] if 200 <= response.status_code < 300: [e2e-llm-inference-service] return response [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Service returned {response.status_code}: {response.text}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] > response = wait_for(get_successful_response, timeout=timeout_seconds, interval=5.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1119: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .get_successful_response at 0x7fbae4adc180> [e2e-llm-inference-service] timeout = 900, interval = 5.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1215: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def get_successful_response(): [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_case.url_getter: [e2e-llm-inference-service] service_url = test_case.url_getter(kserve_client, test_case.llm_service) [e2e-llm-inference-service] else: [e2e-llm-inference-service] service_url = get_llm_service_url(kserve_client, test_case.llm_service) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] raise AssertionError(f"❌ Failed to get service URL: {e}") from e [e2e-llm-inference-service] [e2e-llm-inference-service] model_url = service_url + test_case.endpoint [e2e-llm-inference-service] [e2e-llm-inference-service] headers = {"Content-Type": "application/json"} [e2e-llm-inference-service] if extra_headers: [e2e-llm-inference-service] headers.update(extra_headers) [e2e-llm-inference-service] [e2e-llm-inference-service] if test_case.payload_formatter is not None: [e2e-llm-inference-service] test_payload = test_case.payload_formatter(test_case) [e2e-llm-inference-service] elif test_case.prompt is not None: [e2e-llm-inference-service] test_payload = { [e2e-llm-inference-service] "model": test_case.model_name [e2e-llm-inference-service] if not extra_headers or MODEL_ROUTING_HEADER not in extra_headers [e2e-llm-inference-service] else extra_headers[MODEL_ROUTING_HEADER], [e2e-llm-inference-service] "prompt": test_case.prompt, [e2e-llm-inference-service] "max_tokens": test_case.max_tokens, [e2e-llm-inference-service] } [e2e-llm-inference-service] else: [e2e-llm-inference-service] test_payload = None [e2e-llm-inference-service] [e2e-llm-inference-service] logger.info(f"Calling LLM service at {model_url} with payload {test_payload}") [e2e-llm-inference-service] try: [e2e-llm-inference-service] if test_payload is not None: [e2e-llm-inference-service] response = post_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] json_data=test_payload, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] else: [e2e-llm-inference-service] response = get_with_retry( [e2e-llm-inference-service] model_url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] timeout=test_case.response_timeout, [e2e-llm-inference-service] ) [e2e-llm-inference-service] except Exception as e: [e2e-llm-inference-service] logger.error(f"❌ Failed to call model: {e}") [e2e-llm-inference-service] > raise AssertionError(f"❌ Failed to call model: {e}") from e [e2e-llm-inference-service] E AssertionError: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1109: AssertionError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1681 Created ConfigMap scheduler-config-e2e in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig router-managed-scheduler-config-305f7a8b in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig router-managed-scheduler-config-305f7a8b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig router-managed-scheduler-config-305f7a8b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig scheduler-with-configmap-ref-sc-67492bcd in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig scheduler-with-configmap-ref-sc-67492bcd [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig scheduler-with-configmap-ref-sc-67492bcd [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig workload-llmd-simulator-schedul-b2f159da in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig workload-llmd-simulator-schedul-b2f159da [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig workload-llmd-simulator-schedul-b2f159da [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_inference_service] [2026-07-02T14:13:43.873210] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'scheduler-with-configmap-ref', 'workload-llmd-simulator'], prompt='KServe is a', service_name='scheduler-configmap-ref-test', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': None, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'scheduler-configmap-ref-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-scheduler-config-305f7a8b'}, [e2e-llm-inference-service] {'name': 'scheduler-with-configmap-ref-sc-67492bcd'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-schedul-b2f159da'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-02T14:13:43.886524] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'scheduler-configmap-ref-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-scheduler-config-305f7a8b'}, [e2e-llm-inference-service] {'name': 'scheduler-with-configmap-ref-sc-67492bcd'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-schedul-b2f159da'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-02T14:13:43.998843] end - ✅ in 0.112s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-02T14:13:43.998959] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'scheduler-configmap-ref-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-scheduler-config-305f7a8b'}, [e2e-llm-inference-service] {'name': 'scheduler-with-configmap-ref-sc-67492bcd'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-schedul-b2f159da'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T14:13:57Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T14:13:57Z', 'message': 'Inference Pool kserve-ci-e2e-test/scheduler-configmap-ref-test-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T14:13:57Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:13:57Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T14:13:57Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T14:13:57Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T14:13:57Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:13:57Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T14:14:05Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T14:14:05Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T14:14:05Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:13:57Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T14:14:05Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T14:14:05Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T14:14:05Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:14:05Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T14:14:07Z', 'severity': 'Info', 'status': 'True', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T14:14:07Z', 'severity': 'Info', 'status': 'True', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T14:14:07Z', 'severity': 'Info', 'status': 'True', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:13:57Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T14:14:07Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T14:14:07Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T14:14:07Z', 'message': 'Deployment does not have minimum availability.', 'reason': 'MinimumReplicasUnavailable', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:14:07Z', 'status': 'True', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [wait_for_llm_isvc_ready] [2026-07-02T14:14:28.588955] end - ✅ in 44.590s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_model_response] [2026-07-02T14:14:28.589094] start - args=(, TestCase(base_refs=['router-managed', 'scheduler-with-configmap-ref', 'workload-llmd-simulator'], prompt='KServe is a', service_name='scheduler-configmap-ref-test', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'scheduler-configmap-ref-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-scheduler-config-305f7a8b'}, [e2e-llm-inference-service] {'name': 'scheduler-with-configmap-ref-sc-67492bcd'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-schedul-b2f159da'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m'), 900), kwargs={'extra_headers': None} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [get_llm_service_url] [2026-07-02T14:14:28.589362] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'scheduler-configmap-ref-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-scheduler-config-305f7a8b'}, [e2e-llm-inference-service] {'name': 'scheduler-with-configmap-ref-sc-67492bcd'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-schedul-b2f159da'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [get_llm_service_url] [2026-07-02T14:14:28.598876] end - ✅ in 0.009s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1092 Calling LLM service at http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions with payload {'model': 'facebook/opt-125m', 'prompt': 'KServe is a', 'max_tokens': 20} [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=7, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=6, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=5, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=4, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=3, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=2, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=1, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")': /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=0, connect=None, read=None, redirect=None, status=None)) after connection broken by 'RemoteDisconnected('Remote end closed connection without response')': /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1108 ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:1219 Timed out waiting: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_model_response] [2026-07-02T14:29:33.130999] end - ❌ 904.542s: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:831 [router-managed-scheduler-with-configmap-ref-workload-llmd-simulator] ❌ ERROR: Failed to call llm inference service scheduler-configmap-ref-test: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1240 🔍 # Diagnostics for 'scheduler-configmap-ref-test' in 'kserve-ci-e2e-test' [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1241 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1242 # LLMInferenceService scheduler-configmap-ref-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1245 apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] security.opendatahub.io/enable-auth: 'false' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:43Z' [e2e-llm-inference-service] finalizers: [e2e-llm-inference-service] - serving.kserve.io/llmisvc-finalizer [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:security.opendatahub.io/enable-auth: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:baseRefs: {} [e2e-llm-inference-service] manager: OpenAPI-Generator [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:43Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:finalizers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] v:"serving.kserve.io/llmisvc-finalizer": {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:44Z' [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:addresses: {} [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-decode-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-prefill-worker-data-parallel: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-router-route: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-scheduler: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-template: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-tracing: {} [e2e-llm-inference-service] f:serving.kserve.io/config-llm-worker-data-parallel: {} [e2e-llm-inference-service] f:appliedConfigs: {} [e2e-llm-inference-service] f:conditions: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:router: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:gateways: {} [e2e-llm-inference-service] f:scheduler: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:inferencePool: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:service: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:url: {} [e2e-llm-inference-service] f:workloads: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:primary: {} [e2e-llm-inference-service] f:scheduler: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] resourceVersion: '84499' [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] baseRefs: [e2e-llm-inference-service] - name: router-managed-scheduler-config-305f7a8b [e2e-llm-inference-service] - name: scheduler-with-configmap-ref-sc-67492bcd [e2e-llm-inference-service] - name: workload-llmd-simulator-schedul-b2f159da [e2e-llm-inference-service] model: [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uri: '' [e2e-llm-inference-service] status: [e2e-llm-inference-service] addresses: [e2e-llm-inference-service] - name: gateway-external-model-routing [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/ [e2e-llm-inference-service] - name: gateway-external [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/scheduler-configmap-ref-test [e2e-llm-inference-service] - name: gateway-internal-model-routing [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/ [e2e-llm-inference-service] - name: gateway-internal [e2e-llm-inference-service] url: http://openshift-ai-inference-openshift-default.openshift-ingress.svc.cluster.local/kserve-ci-e2e-test/scheduler-configmap-ref-test [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-template: kserve-config-llm-decode-template [e2e-llm-inference-service] serving.kserve.io/config-llm-decode-worker-data-parallel: kserve-config-llm-decode-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-template: kserve-config-llm-prefill-template [e2e-llm-inference-service] serving.kserve.io/config-llm-prefill-worker-data-parallel: kserve-config-llm-prefill-worker-data-parallel [e2e-llm-inference-service] serving.kserve.io/config-llm-router-route: kserve-config-llm-router-route [e2e-llm-inference-service] serving.kserve.io/config-llm-scheduler: kserve-config-llm-scheduler [e2e-llm-inference-service] serving.kserve.io/config-llm-template: kserve-config-llm-template [e2e-llm-inference-service] serving.kserve.io/config-llm-tracing: kserve-config-llm-tracing [e2e-llm-inference-service] serving.kserve.io/config-llm-worker-data-parallel: kserve-config-llm-worker-data-parallel [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:07Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: HTTPRoutesReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:07Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: InferencePoolReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:07Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: MainWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:13:57Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: PresetsCombined [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Ready [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: RouterReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] severity: Info [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: SchedulerWorkloadReady [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:07Z' [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: WorkloadsReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] url: http://a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com/kserve-ci-e2e-test/scheduler-configmap-ref-test [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:44 TIME NAMESPACE SOURCE TYPE REASON MESSAGE [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:45 -------------------------------------------------------------------------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-6599bc8fd-8sr69 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-6599bc8fd-8sr69 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-router-scheduler-6fcbb55f87 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-6599bc8fd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:52 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy auth-disabled-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "auth-disabled-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-disabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-disabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-disabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-disabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:56 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-disabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-67c57d579b-gtwnj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:05 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 27.649s (27.649s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.35:8000/health": dial tcp 10.132.0.35:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-67c57d579b-gtwnj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.31/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.181s (1.181s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-router-scheduler-688c79755d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-67c57d579b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-enabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-enabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-enabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-enabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-enabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-enabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-6cd6c5884-rrbrr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-6cd6c5884-rrbrr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-router-scheduler-5b958c76d8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-6cd6c5884 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-invalid-token-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-invalid-token-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-invalid-token-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-invalid-token-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-invalid-token-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:24 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-invalid-token-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:25 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.52:8001/health": dial tcp 10.132.0.52:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.53/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.53:8000/health": dial tcp 10.132.0.53:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-router-scheduler-f7968dfdd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-79d6b9cc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:31 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-pd-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-pd-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:43 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-98989556b-vl89r to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:14:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.45:8000/health": dial tcp 10.132.0.45:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-98989556b-vl89r [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-router-scheduler-579b7545d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-98989556b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:57 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:15:01 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/e2e-pvc-model-download-4s522 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:26 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test job-controller Normal SuccessfulCreate Created pod: e2e-pvc-model-download-4s522 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:36 kserve-ci-e2e-test job-controller Normal Completed Job completed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:20 kserve-ci-e2e-test persistentvolume-controller Normal WaitForFirstConsumer waiting for first consumer to be created before binding [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test persistentvolume-controller Normal ExternalProvisioning Waiting for a volume to be created either by the external provisioner 'ebs.csi.aws.com' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal Provisioning External provisioner is provisioning volume for claim "kserve-ci-e2e-test/e2e-pvc-model-storage" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal ProvisioningSucceeded Successfully provisioned volume pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.094s (1.094s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:28 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec0c69dceeb48768325d1a53a749e65786-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.30/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedMount MountVolume.SetUp failed for volume "tls-certs" : secret "gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.60/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:46 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.60:8000/health": dial tcp 10.132.0.60:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.60:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:32 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-3c960099-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-3c960099-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:02 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-3c960099] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" in 808ms (808ms including waiting). Image size: 44914394 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.48:8001/health": dial tcp 10.132.0.48:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.49:8000/health": dial tcp 10.132.0.49:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5b86594bc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:44 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:52 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-50bc673d] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:07:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.44:8000/health": dial tcp 10.132.0.44:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:07:54 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-87882a8e] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 44.193s (44.193s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.62/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:40:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.62:8000/health": dial tcp 10.132.0.62:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 7f6c8cfc6 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:40:17 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test386bd5808c5c450f11fd8632e1f821fb-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:58 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-dc21cb14] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:20 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.45:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:27 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-e95b1dc1] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.42:8000/health": dial tcp 10.132.0.42:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler-67f9b884d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:01 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:19 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-7ca60146] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:44 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler-5444b4dbf from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:57 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:55 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-ba4d693a] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.55/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.54/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:25 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 686d468674 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:35 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-2577e794-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:53 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-2577e794] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:56 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.47:8000/health": dial tcp 10.132.0.47:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:51 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-59b9d263-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-59b9d263-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:25 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-59b9d263] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.50:8001/health": dial tcp 10.132.0.50:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.51:8000/health": dial tcp 10.132.0.51:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-79497db4cc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:56 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-e8706282-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:09 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-e8706282] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test80f99418f0f3eb4c80f02e6144a6d4f2-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:36 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4c1ba212] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:30 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.34:8000/health": dial tcp 10.133.0.34:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:36 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:14 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4f8c0978] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.32/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:50 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:48 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-a50492e9] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:40 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-4b931143-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-4b931143-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:21 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-4b931143] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:25 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.42:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-5b1e8f15-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-5b1e8f15-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:22 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-5b1e8f15] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-e45d1f79-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-e45d1f79-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:20 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-e45d1f79] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler-6fcb489785 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.566s (1.566s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-7dbcb75dbc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0, llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler-79866dc455 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:03 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedToRetrieveImagePullSecret Unable to retrieve some image pull secrets (llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa-dockercfg-pg2h6); attempting to pull the image may not succeed. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc44d181485fad85e662eb092f3749502f-kserve-router-scheduler-69bf65579d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler-5dd88bfbb7 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler-5f555d4d85 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler-54ccf64f4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler-6d86bd4d9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-67b4bb9646 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9, llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler-68cc9685d6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler-749449dbc8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-ccbb969bd-65ln7 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.63/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:56:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.63:8000/health": dial tcp 10.132.0.63:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-multiple-adapters-test-kserve-ccbb969bd-65ln7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-multiple-adapters-test-kserve-ccbb969bd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:08 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-multiple-adapters-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-multiple-adapters-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-multiple-adapters-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:56:42 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-multiple-adapters-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.61/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-single-adapter-test-kserve-85cd6c88dc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-single-adapter-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-single-adapter-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-single-adapter-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-single-adapter-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.29/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.483s (1.483s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.29:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 3.651s (3.651s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.88s (1.88s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 2.962s (2.962s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.723s (1.723s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" in 30.447s (30.447s including waiting). Image size: 2989890188 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Liveness probe failed: timeout: failed to connect service "10.132.0.34:9003" within 1s: context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-router-scheduler-77d88bdcf4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-7c94f9d7c6 from 0 to 2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy precise-prefix-cache-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "precise-prefix-cache-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/precise-prefix-cache-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/precise-prefix-cache-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/precise-prefix-cache-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/precise-prefix-cache-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:15 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [precise-prefix-cache-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:16 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/prestop-hook-test-kserve-66474c4c54-j59xs to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.67/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:29 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.67:8000/health": dial tcp 10.132.0.67:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: prestop-hook-test-kserve-66474c4c54-j59xs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/prestop-hook-test-kserve-router-scheduler-7c7698955f-mvsc8 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:00 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.68/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: prestop-hook-test-kserve-router-scheduler-7c7698955f-mvsc8 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set prestop-hook-test-kserve-router-scheduler-7c7698955f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set prestop-hook-test-kserve-66474c4c54 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/prestop-hook-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/prestop-hook-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/prestop-hook-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/prestop-hook-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-prestop-hook-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/prestop-hook-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/prestop-hook-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/prestop-hook-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/prestop-hook-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/prestop-hook-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/prestop-hook-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/prestop-hook-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-1-openshift-default-799f46c59b-fnbwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.28/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:31 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "http://10.133.0.28:15021/healthz/ready": context deadline exceeded (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning BackOff Back-off restarting failed container istio-proxy in pod router-gateway-1-openshift-default-799f46c59b-fnbwc_kserve-ci-e2e-test(f244ca92-1cf1-429e-886d-e0d585d0175f) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.28:15021/healthz/ready": dial tcp 10.133.0.28:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-1-openshift-default-799f46c59b-fnbwc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:58 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-1-openshift-default-799f46c59b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:37 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-2-openshift-default-54c789bdc6-w57g9 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:19 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.44:15021/healthz/ready": dial tcp 10.133.0.44:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-2-openshift-default-54c789bdc6-w57g9 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-2-openshift-default-54c789bdc6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:17 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.56:8001/health": dial tcp 10.132.0.56:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.57/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.57:8000/health": dial tcp 10.132.0.57:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-prefill-778968b98b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-router-scheduler-5c684fd54d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-694c6ddf4c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:33 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:47 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-587fcd8566-2pqvl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:20:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:29 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-587fcd8566-2pqvl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-router-scheduler-74545d8489 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-587fcd8566 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:05 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.66/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-configmap-ref-test-kserve-router-scheduler-68f876cc5d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-configmap-ref-test-kserve-744bf59fbd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy scheduler-configmap-ref-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "scheduler-configmap-ref-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-configmap-ref-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [scheduler-configmap-ref-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-7c6c58cc96-l62qh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-7c6c58cc96-l62qh [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-router-scheduler-d9ddf6c4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-7c6c58cc96 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy scheduler-inline-config-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "scheduler-inline-config-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/scheduler-inline-config-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/scheduler-inline-config-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/scheduler-inline-config-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/scheduler-inline-config-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [scheduler-inline-config-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-6hzlt to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.59/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:51 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-f2dmd to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.58/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-f2dmd [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-6hzlt [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:07 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy stop-feature-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "stop-feature-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-stop-feature-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/stop-feature-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/stop-feature-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/stop-feature-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [stop-feature-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.InferencePool kserve-ci-e2e-test/stop-feature-test-inference-pool [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:56 kserve-ci-e2e-test LLMInferenceServiceController Warning LLMInferenceServiceNotReady LLMInferenceService [stop-feature-test] is no longer Ready because of: GatewaysReady, HTTPRoutesReady, InferencePoolReady, MainWorkloadReady, PrefillWorkerWorkloadReady, PrefillWorkloadReady, RouterReady, SchedulerWorkloadReady, WorkerWorkloadReady, WorkloadsReady [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/tls-verification-test-kserve-6d6946dcd6-bdddw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.64/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.64:8000/health": dial tcp 10.132.0.64:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.64:8000/health": dial tcp 10.132.0.64:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: tls-verification-test-kserve-6d6946dcd6-bdddw [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.65/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set tls-verification-test-kserve-router-scheduler-7f955879f5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set tls-verification-test-kserve-6d6946dcd6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:24 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy tls-verification-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/tls-verification-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "tls-verification-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/tls-verification-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/tls-verification-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-tls-verification-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/tls-verification-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/tls-verification-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/tls-verification-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/tls-verification-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/tls-verification-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [tls-verification-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-tls-verification-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 I0702 14:13:55.362585 1 config.go:602] "Configuration:" =< [e2e-llm-inference-service] { [e2e-llm-inference-service] "IP": "", [e2e-llm-inference-service] "PodName": "", [e2e-llm-inference-service] "PodNameSpace": "", [e2e-llm-inference-service] "VllmDevMode": false, [e2e-llm-inference-service] "block-size": 16, [e2e-llm-inference-service] "data-parallel-rank": -1, [e2e-llm-inference-service] "data-parallel-size": 1, [e2e-llm-inference-service] "dataset-in-memory": false, [e2e-llm-inference-service] "dataset-path": "", [e2e-llm-inference-service] "dataset-table-name": "llmd", [e2e-llm-inference-service] "dataset-url": "", [e2e-llm-inference-service] "default-embedding-dimensions": 384, [e2e-llm-inference-service] "ec-transfer-config": "", [e2e-llm-inference-service] "enable-kvcache": false, [e2e-llm-inference-service] "enable-prefix-caching": false, [e2e-llm-inference-service] "enable-request-id-headers": false, [e2e-llm-inference-service] "enable-sleep-mode": false, [e2e-llm-inference-service] "enforce-eager": false, [e2e-llm-inference-service] "event-batch-size": 16, [e2e-llm-inference-service] "failure-injection-rate": 0, [e2e-llm-inference-service] "failure-types": null, [e2e-llm-inference-service] "fake-metrics": null, [e2e-llm-inference-service] "fake-metrics-refresh-interval": 100000000, [e2e-llm-inference-service] "global-cache-hit-threshold": 0, [e2e-llm-inference-service] "hash-seed": "", [e2e-llm-inference-service] "inter-token-latency": 0, [e2e-llm-inference-service] "inter-token-latency-std-dev": 0, [e2e-llm-inference-service] "kv-cache-size": 1024, [e2e-llm-inference-service] "kv-cache-transfer-latency": 0, [e2e-llm-inference-service] "kv-cache-transfer-latency-std-dev": 0, [e2e-llm-inference-service] "kv-cache-transfer-time-per-token": 0, [e2e-llm-inference-service] "kv-cache-transfer-time-std-dev": 0, [e2e-llm-inference-service] "latency-calculator": "", [e2e-llm-inference-service] "lora-modules": null, [e2e-llm-inference-service] "max-cpu-loras": 1, [e2e-llm-inference-service] "max-loras": 1, [e2e-llm-inference-service] "max-model-len": 1024, [e2e-llm-inference-service] "max-num-seqs": 5, [e2e-llm-inference-service] "max-tool-call-array-param-length": 5, [e2e-llm-inference-service] "max-tool-call-integer-param": 100, [e2e-llm-inference-service] "max-tool-call-number-param": 100, [e2e-llm-inference-service] "max-waiting-queue-length": 1000, [e2e-llm-inference-service] "min-tool-call-array-param-length": 1, [e2e-llm-inference-service] "min-tool-call-integer-param": 0, [e2e-llm-inference-service] "min-tool-call-number-param": 0, [e2e-llm-inference-service] "mm-encoder-only": false, [e2e-llm-inference-service] "mm-processor-kwargs": "", [e2e-llm-inference-service] "mode": "random", [e2e-llm-inference-service] "model": "facebook/opt-125m", [e2e-llm-inference-service] "object-tool-call-not-required-field-probability": 50, [e2e-llm-inference-service] "port": 8000, [e2e-llm-inference-service] "prefill-overhead": 0, [e2e-llm-inference-service] "prefill-time-per-token": 0, [e2e-llm-inference-service] "prefill-time-std-dev": 0, [e2e-llm-inference-service] "seed": 1783001635362180000, [e2e-llm-inference-service] "self-signed-certs": false, [e2e-llm-inference-service] "served-model-name": [ [e2e-llm-inference-service] "facebook/opt-125m" [e2e-llm-inference-service] ], [e2e-llm-inference-service] "ssl-certfile": "/var/run/kserve/tls/tls.crt", [e2e-llm-inference-service] "ssl-keyfile": "/var/run/kserve/tls/tls.key", [e2e-llm-inference-service] "time-factor-under-load": 1, [e2e-llm-inference-service] "time-to-first-token": 0, [e2e-llm-inference-service] "time-to-first-token-std-dev": 0, [e2e-llm-inference-service] "tool-call-not-required-param-probability": 50, [e2e-llm-inference-service] "uds-socket-path": "/tmp/tokenizer/tokenizer-uds.socket", [e2e-llm-inference-service] "zmq-endpoint": "tcp://127.0.0.1:5557" [e2e-llm-inference-service] } [e2e-llm-inference-service] > [e2e-llm-inference-service] I0702 14:13:55.405961 1 tokenizer.go:104] "Model is not a real HF model, using simulated tokenizer" model="facebook/opt-125m" [e2e-llm-inference-service] I0702 14:13:55.410059 1 context.go:138] "No dataset path or URL provided, using random text for responses" [e2e-llm-inference-service] I0702 14:13:55.410117 1 communication.go:49] "Starting communication layer" [e2e-llm-inference-service] I0702 14:13:55.410629 1 http_server_tls.go:44] "HTTPS server starting with certificate files" cert="/var/run/kserve/tls/tls.crt" key="/var/run/kserve/tls/tls.key" [e2e-llm-inference-service] I0702 14:13:55.410642 1 simulator.go:188] "Start processing routine" [e2e-llm-inference-service] I0702 14:13:55.410721 1 grpc.go:126] "Server starting" protocol="gRPC" port=8000 [e2e-llm-inference-service] I0702 14:13:55.411296 1 http.go:96] "Server starting" protocol="HTTPS" port=8000 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:201 # -- logs (current) -- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:202 {"level":"info","ts":1783001635.5996876,"logger":"setup","caller":"runner/runner.go:196","msg":"GIE build","commit-sha":"181aa8358916e19b8844ccc752b2d6153d4b2ad6","build-ref":"v0.9.0-rc.2"} [e2e-llm-inference-service] Flag --model-server-metrics-scheme has been deprecated, This flag is deprecated. Configure via EndpointPickerConfig data layer plugin parameters instead. [e2e-llm-inference-service] {"level":"info","ts":1783001635.5998049,"logger":"setup","caller":"runner/runner.go:217","msg":"Flags processed","flags":{"cert-path":"/var/run/kserve/tls","config-file":"","config-text":"apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\nplugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type: prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n- name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n","disable-endpoint-subset-filter":false,"enable-cert-reload":true,"enable-grpc-stream-metrics":false,"enable-pprof":true,"endpoint-selector":"","endpoint-target-ports":{},"grpc-health-port":9003,"grpc-max-recv-msg-size":"","grpc-max-send-msg-size":"","grpc-port":9002,"ha-enable-leader-election":false,"health-checking":false,"metrics-endpoint-auth":true,"metrics-port":9090,"metrics-staleness-threshold":2000000000,"model-server-metrics-https-insecure-skip-verify":true,"model-server-metrics-path":"/metrics","model-server-metrics-port":0,"model-server-metrics-scheme":"https","pool-group":"inference.networking.k8s.io","pool-name":"scheduler-configmap-ref-test-inference-pool","pool-namespace":"kserve-ci-e2e-test","refresh-metrics-interval":50000000,"refresh-prometheus-metrics-interval":5000000000,"secure-serving":true,"tracing":true,"v":2,"zap-devel":{},"zap-encoder":{},"zap-log-level":{},"zap-stacktrace-level":{},"zap-time-encoding":{}}} [e2e-llm-inference-service] {"level":"info","ts":1783001635.5998974,"logger":"setup.trace","caller":"tracing/telemetry.go:123","msg":"init OTel trace exporter","type":"console"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6003451,"caller":"loader/configloader.go:89","msg":"DEPRECATION: apiVersion inference.networking.x-k8s.io/v1alpha1/EndpointPickerConfig is deprecated","replacement":"llm-d.ai/v1alpha1/EndpointPickerConfig"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6003792,"caller":"loader/configloader.go:121","msg":"Loaded raw configuration","config":"{Plugins: [{Type: single-profile-handler} {Type: queue-scorer} {Type: prefix-cache-scorer} {Type: max-score-picker}], SchedulingProfiles: [{Name: default, Plugins: [{PluginRef: queue-scorer, Weight: 2.00} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: max-score-picker}]}]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6003902,"logger":"setup","caller":"runner/runner.go:622","msg":"Data layer: ENABLED"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6006582,"logger":"setup","caller":"runner/runner.go:281","msg":"Raw config after phase one","config":{"apiVersion":"inference.networking.x-k8s.io/v1alpha1","dataLayer":null,"kind":"EndpointPickerConfig","plugins":[{"name":"single-profile-handler","parameters":null,"type":"single-profile-handler"},{"name":"queue-scorer","parameters":null,"type":"queue-scorer"},{"name":"prefix-cache-scorer","parameters":null,"type":"prefix-cache-scorer"},{"name":"max-score-picker","parameters":null,"type":"max-score-picker"}],"schedulingProfiles":[{"name":"default","plugins":[{"pluginRef":"queue-scorer","weight":2},{"pluginRef":"prefix-cache-scorer","weight":3},{"pluginRef":"max-score-picker","weight":null}]}]}} [e2e-llm-inference-service] {"level":"info","ts":1783001635.621277,"logger":"utilization-detector/utilization-detector","caller":"utilization/detector.go:83","msg":"Creating new UtilizationDetector","queueDepthThreshold":5,"kvCacheUtilThreshold":0.8,"metricsStalenessThreshold":"200ms","headroom":0} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6213968,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"vllm","mapping":"Mapping{all specs enabled}"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6214318,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"sglang","mapping":"Mapping{disabled: [lora]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6214688,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"trtllm-serve","mapping":"Mapping{disabled: [lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.621539,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"triton-tensorrt-llm","mapping":"Mapping{disabled: [lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6215584,"caller":"metrics/factories.go:230","msg":"Registered engine mapping","engine":"triton","mapping":"Mapping{disabled: [kv, lora, cacheInfo]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6216078,"caller":"loader/configloader.go:154","msg":"Instantiated all plugins and applied system defaults. Effective raw configuration","config":"{Plugins: [{Name: single-profile-handler, Type: single-profile-handler} {Name: queue-scorer, Type: queue-scorer} {Name: prefix-cache-scorer, Type: prefix-cache-scorer} {Name: max-score-picker, Type: max-score-picker} {Name: fcfs-ordering-policy, Type: fcfs-ordering-policy} {Name: global-strict-fairness-policy, Type: global-strict-fairness-policy} {Name: static-usage-limit-policy, Type: static-usage-limit-policy} {Name: openai-parser, Type: openai-parser} {Name: anthropic-parser, Type: anthropic-parser} {Name: vllmhttp-parser, Type: vllmhttp-parser} {Name: utilization-detector, Type: utilization-detector} {Name: metrics-data-source, Type: metrics-data-source} {Name: core-metrics-extractor, Type: core-metrics-extractor}], SchedulingProfiles: [{Name: default, Plugins: [{PluginRef: queue-scorer, Weight: 2.00} {PluginRef: prefix-cache-scorer, Weight: 3.00} {PluginRef: max-score-picker}]}], DataLayer: {Sources: [{PluginRef: metrics-data-source, Extractors: [{PluginRef: core-metrics-extractor}]}], Discovery: }, FlowControl: {MaxBytes: unlimited, MaxRequests: unlimited, SaturationDetector: {PluginRef: utilization-detector}}, RequestHandler: {Parsers: [{PluginRef: openai-parser}, {PluginRef: anthropic-parser}, {PluginRef: vllmhttp-parser}]}}"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6216562,"caller":"approximateprefix/plugin.go:88","msg":"Prefix DataProducer initialized","config":{"autoTune":true,"blockSizeTokens":16,"blockSize":0,"maxPrefixBlocksToMatch":2048,"maxPrefixTokensToMatch":131072,"lruCapacityPerServer":31250}} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6217303,"caller":"approximateprefix/plugin.go:111","msg":"WARNING: configured blockSizeTokens is below the recommended minimum, overriding it.","blockSizeTokens":16,"minimum":64,"issue":"https://github.com/llm-d/llm-d-router/issues/1158"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6217587,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"approx-prefix-cache-producer/approx-prefix-cache-producer","dataKey":"PrefixCacheMatchInfoDataKey/approx-prefix-cache-producer","consumer":"prefix-cache-scorer"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.621802,"caller":"datalayer/data_graph.go:116","msg":"auto-created default producer","producer":"token-producer/token-producer","dataKey":"TokenizedPrompt/token-producer","consumer":"approx-prefix-cache-producer"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.621879,"caller":"runner/runner.go:685","msg":"loaded configuration from file/text successfully"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6218894,"logger":"setup","caller":"runner/runner.go:308","msg":"EPP config after phase two","config":"{SchedulerConfig:{ProfileHandler: single-profile-handler/single-profile-handler, Profiles: map[default:{Filters: [], Scorers: [queue-scorer/queue-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000], Picker: max-score-picker/max-score-picker}]} SaturationDetector:0xc0008349c0 DataConfig:{Sources:[{Plugin:0xc0004bd4d0 Extractors:[0xc000834cc0]}]} FlowControlConfig: ParserRegistry:0xc000835280}"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6403797,"logger":"setup","caller":"runner/runner.go:352","msg":"Setting pprof handlers"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404083,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/profile"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404214,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/trace"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404262,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/heap"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404307,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/allocs"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404347,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/threadcreate"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.640439,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/mutex"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404433,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404471,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/symbol"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404555,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/goroutine"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.64046,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/block"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404643,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/pprof/cmdline"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404757,"caller":"manager/internal.go:201","msg":"Registering metrics http server extra handler","path":"/debug/plugins/state"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6404834,"logger":"setup","caller":"runner/runner.go:373","msg":"parsed config","scheduler-config":"{ProfileHandler: single-profile-handler/single-profile-handler, Profiles: map[default:{Filters: [], Scorers: [queue-scorer/queue-scorer: 2.000000, prefix-cache-scorer/prefix-cache-scorer: 3.000000], Picker: max-score-picker/max-score-picker}]}"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6405075,"logger":"setup","caller":"datalayer/runtime.go:99","msg":"Configuring datalayer runtime","numSources":1} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6405158,"logger":"setup","caller":"datalayer/runtime.go:118","msg":"Processing source","source":"metrics-data-source","numExtractors":1} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6405292,"logger":"setup","caller":"datalayer/runtime.go:147","msg":"Source configured","source":"metrics-data-source","extractors":["core-metrics-extractor/core-metrics-extractor"]} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6405427,"logger":"setup","caller":"datalayer/runtime.go:206","msg":"Datalayer runtime configured","pollers":1,"notifiers":0,"endpointSources":0} [e2e-llm-inference-service] {"level":"info","ts":1783001635.640552,"logger":"setup","caller":"runner/runner.go:833","msg":"Experimental Flow Control layer is disabled, using legacy admission control"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.64064,"logger":"setup","caller":"runner/runner.go:721","msg":"ExtProc server runner added to manager."} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6406531,"logger":"setup","caller":"runner/runner.go:260","msg":"Controller manager starting"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6406772,"logger":"controller-runtime.metrics","caller":"server/server.go:208","msg":"Starting metrics server"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6408505,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"health"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6412024,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","source":"kind source: *v1.InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6412194,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite","source":"kind source: *v1alpha2.InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6413198,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"health","port":9003} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6413803,"logger":"controller-runtime.metrics","caller":"server/server.go:247","msg":"Serving metrics server","bindAddress":":9090","secure":false} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6418576,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"pod","controllerGroup":"","controllerKind":"Pod","source":"kind source: *v1.Pod"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.641951,"caller":"controller/controller.go:370","msg":"Starting EventSource","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective","source":"kind source: *v1alpha2.InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6421034,"caller":"runnable/grpc.go:35","msg":"gRPC server starting","name":"ext-proc"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6422029,"caller":"runnable/grpc.go:43","msg":"gRPC server listening","name":"ext-proc","port":9002} [e2e-llm-inference-service] {"level":"info","ts":1783001635.645806,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceModelRewrite","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6461563,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1alpha2.InferenceObjective","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6462462,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.InferencePool","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.6481323,"logger":"controller-runtime.cache","caller":"cache/reflector.go:446","msg":"Caches populated","type":"*v1.Pod","reflector":"pkg/mod/k8s.io/client-go@v0.35.6/tools/cache/reflector.go:289"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.742288,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.7423146,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencemodelrewrite","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceModelRewrite","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783001635.742283,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.742342,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783001635.7424872,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"scheduler-configmap-ref-test-inference-pool","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"scheduler-configmap-ref-test-inference-pool","reconcileID":"d889c603-456a-4a42-98b7-56418d4b0fe9"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.7428489,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.7428613,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"inferenceobjective","controllerGroup":"inference.networking.x-k8s.io","controllerKind":"InferenceObjective","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783001635.8436193,"caller":"controller/controller.go:303","msg":"Starting Controller","controller":"pod","controllerGroup":"","controllerKind":"Pod"} [e2e-llm-inference-service] {"level":"info","ts":1783001635.8436465,"caller":"controller/controller.go:306","msg":"Starting workers","controller":"pod","controllerGroup":"","controllerKind":"Pod","worker count":1} [e2e-llm-inference-service] {"level":"info","ts":1783001643.5687282,"caller":"controller/inferencepool_reconciler.go:46","msg":"Reconciling InferencePool","controller":"inferencepool","controllerGroup":"inference.networking.k8s.io","controllerKind":"InferencePool","InferencePool":{"name":"scheduler-configmap-ref-test-inference-pool","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"scheduler-configmap-ref-test-inference-pool","reconcileID":"10264311-53e4-4a5c-a694-b1fd012232aa"} [e2e-llm-inference-service] {"level":"info","ts":1783001646.1718304,"caller":"controller/pod_reconciler.go:99","msg":"Pod already exists","controller":"pod","controllerGroup":"","controllerKind":"Pod","Pod":{"name":"scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk","namespace":"kserve-ci-e2e-test"},"namespace":"kserve-ci-e2e-test","name":"scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk","reconcileID":"5ff40b61-2efb-4444-8c4d-909555319230"} [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-service [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 62d9f9ed-2569-4c34-9813-c87e5416f376 [e2e-llm-inference-service] resourceVersion: '84492' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.132.0.66 [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr [e2e-llm-inference-service] uid: cbb15ca7-fd08-4aba-9cda-b969e72cdad8 [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 5253c225-5b85-4e1f-b7d4-da32e7457dd8 [e2e-llm-inference-service] resourceVersion: '84314' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - ip: 10.134.0.56 [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk [e2e-llm-inference-service] uid: a9963664-c152-42c0-b9dc-bed11b2e32d2 [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk [e2e-llm-inference-service] generateName: scheduler-configmap-ref-test-kserve-744bf59fbd- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: a9963664-c152-42c0-b9dc-bed11b2e32d2 [e2e-llm-inference-service] resourceVersion: '84312' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 744bf59fbd [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.134.0.56/23"],"mac_address":"0a:58:0a:86:00:38","gateway_ips":["10.134.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.134.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.134.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.134.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.134.0.1"}],"ip_address":"10.134.0.56/23","gateway_ip":"10.134.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.134.0.56\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:86:00:38\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-744bf59fbd [e2e-llm-inference-service] uid: 474e838e-3aee-4e12-a261-808ecac22cf5 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-138-159 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"474e838e-3aee-4e12-a261-808ecac22cf5"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.134.0.56"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kube-api-access-2pcqv [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/llm-d-inference-sim [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --model [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --mode [e2e-llm-inference-service] - random [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kube-api-access-2pcqv [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: default [e2e-llm-inference-service] serviceAccount: default [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] fsGroup: 1000690000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:56Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] hostIP: 10.0.138.159 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.138.159 [e2e-llm-inference-service] podIP: 10.134.0.56 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.134.0.56 [e2e-llm-inference-service] startTime: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] imageID: ghcr.io/llm-d/llm-d-inference-sim@sha256:bab162bd25e2ed8b15022387cdb223023aeb33be49476af9f0115c0398fb8ff5 [e2e-llm-inference-service] containerID: cri-o://a4f6dc07c521efb90bc4d01871f407a5733e7e7b33371cb72d19de55a3ba1f7b [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-2pcqv [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr [e2e-llm-inference-service] generateName: scheduler-configmap-ref-test-kserve-router-scheduler-68f876cc5d- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: cbb15ca7-fd08-4aba-9cda-b969e72cdad8 [e2e-llm-inference-service] resourceVersion: '84491' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 68f876cc5d [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.132.0.66/23"],"mac_address":"0a:58:0a:84:00:42","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.66/23","gateway_ip":"10.132.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.132.0.66\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:84:00:42\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-router-scheduler-68f876cc5d [e2e-llm-inference-service] uid: a0f37028-ba7d-41fb-a49c-4b3b3fed2ed1 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-134-1 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"a0f37028-ba7d-41fb-a49c-4b3b3fed2ed1"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.132.0.66"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kube-api-access-pk98z [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type: prefix-cache-scorer\n\ [e2e-llm-inference-service] - type: max-score-picker\nschedulingProfiles:\n- name: default\n plugins:\n\ [e2e-llm-inference-service] \ - pluginRef: queue-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n\ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-pk98z [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] serviceAccount: scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] fsGroup: 1000690000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: scheduler-configmap-ref-test-epp-sa-dockercfg-gh7mk [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:56Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] hostIP: 10.0.134.1 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.134.1 [e2e-llm-inference-service] podIP: 10.132.0.66 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.132.0.66 [e2e-llm-inference-service] startTime: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] imageID: ghcr.io/llm-d/llm-d-router-endpoint-picker@sha256:06b6c75d77afd0e07053402752a9736c2dfbc12a306d0d37d963aac4c1d4e6a6 [e2e-llm-inference-service] containerID: cri-o://4a6a7e2598848cd6f7df9914f48ecdff11836701d72778dcc7358b1b33e43167 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-pk98z [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 09423b58-8cbb-4557-8f50-964a94a5baf5 [e2e-llm-inference-service] resourceVersion: '84071' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: scheduler-configmap-ref-test-epp-sa-dockercfg-gh7mk [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"scheduler-configmap-ref-test-epp-sa-dockercfg-gh7mk"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-x9l7t"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: scheduler-configmap-ref-test-epp-sa-dockercfg-gh7mk [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: scheduler-configmap-ref-test-epp-sa-dockercfg-gh7mk [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-service [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 9184a09b-fe50-4859-ab21-51fadf9dc1f8 [e2e-llm-inference-service] resourceVersion: '84090' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] targetPort: grpc [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] targetPort: grpc-health [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] targetPort: metrics [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] targetPort: zmq [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] clusterIP: 172.31.55.156 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.55.156 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 802289f1-228c-4928-aa49-dcefa3a51844 [e2e-llm-inference-service] resourceVersion: '84059' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] targetPort: 8000 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] clusterIP: 172.31.88.245 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.88.245 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 4d3021e7-276e-4d8a-95bc-4371e302d4d3 [e2e-llm-inference-service] resourceVersion: '84318' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:rollingUpdate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:maxSurge: {} [e2e-llm-inference-service] f:maxUnavailable: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/llm-d-inference-sim [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --model [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --mode [e2e-llm-inference-service] - random [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: RollingUpdate [e2e-llm-inference-service] rollingUpdate: [e2e-llm-inference-service] maxUnavailable: 25% [e2e-llm-inference-service] maxSurge: 25% [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "scheduler-configmap-ref-test-kserve-744bf59fbd" has successfully [e2e-llm-inference-service] progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-router-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 9891ccd9-490f-4077-9ce1-ddabfc4b5e1f [e2e-llm-inference-service] resourceVersion: '84495' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n-\ [e2e-llm-inference-service] \ name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] serviceAccount: scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: Recreate [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] reason: MinimumReplicasAvailable [e2e-llm-inference-service] message: Deployment has minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] reason: NewReplicaSetAvailable [e2e-llm-inference-service] message: ReplicaSet "scheduler-configmap-ref-test-kserve-router-scheduler-68f876cc5d" [e2e-llm-inference-service] has successfully progressed. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-744bf59fbd [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 474e838e-3aee-4e12-a261-808ecac22cf5 [e2e-llm-inference-service] resourceVersion: '84317' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 744bf59fbd [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '2' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve [e2e-llm-inference-service] uid: 4d3021e7-276e-4d8a-95bc-4371e302d4d3 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"4d3021e7-276e-4d8a-95bc-4371e302d4d3"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 744bf59fbd [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 744bf59fbd [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/llm-d-inference-sim [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --model [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --mode [e2e-llm-inference-service] - random [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-router-scheduler-68f876cc5d [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: a0f37028-ba7d-41fb-a49c-4b3b3fed2ed1 [e2e-llm-inference-service] resourceVersion: '84494' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 68f876cc5d [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-router-scheduler [e2e-llm-inference-service] uid: 9891ccd9-490f-4077-9ce1-ddabfc4b5e1f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"9891ccd9-490f-4077-9ce1-ddabfc4b5e1f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:availableReplicas: {} [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:readyReplicas: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 68f876cc5d [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 68f876cc5d [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n-\ [e2e-llm-inference-service] \ name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] serviceAccount: scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] readyReplicas: 1 [e2e-llm-inference-service] availableReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-rb [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 146b4813-9bc7-4b31-a45b-fd46f61c36e2 [e2e-llm-inference-service] resourceVersion: '84084' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-role [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-role [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 1d0cbec6-2f02-4957-b9e1-1dc999d18ede [e2e-llm-inference-service] resourceVersion: '84082' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] - create [e2e-llm-inference-service] - update [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - delete [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-service-str2k [e2e-llm-inference-service] generateName: scheduler-configmap-ref-test-epp-service- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 765ed193-8182-4bd4-acf7-b01733d0875e [e2e-llm-inference-service] resourceVersion: '84493' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: scheduler-configmap-ref-test-epp-service [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-service [e2e-llm-inference-service] uid: 9184a09b-fe50-4859-ab21-51fadf9dc1f8 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:14:27Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"9184a09b-fe50-4859-ab21-51fadf9dc1f8"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.132.0.66 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr [e2e-llm-inference-service] uid: cbb15ca7-fd08-4aba-9cda-b969e72cdad8 [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc-njbt9 [e2e-llm-inference-service] generateName: scheduler-configmap-ref-test-kserve-workload-svc- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 05d21d64-3bd5-4463-b8d4-7f8f4fb3cdae [e2e-llm-inference-service] resourceVersion: '84315' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] uid: 802289f1-228c-4928-aa49-dcefa3a51844 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:14:06Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"802289f1-228c-4928-aa49-dcefa3a51844"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.134.0.56 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: true [e2e-llm-inference-service] serving: true [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk [e2e-llm-inference-service] uid: a9963664-c152-42c0-b9dc-bed11b2e32d2 [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-rb [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 146b4813-9bc7-4b31-a45b-fd46f61c36e2 [e2e-llm-inference-service] resourceVersion: '84084' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:kserve-ci-e2e-test:scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-role [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-role [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 1d0cbec6-2f02-4957-b9e1-1dc999d18ede [e2e-llm-inference-service] resourceVersion: '84082' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:13:54Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - create [e2e-llm-inference-service] - delete [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - update [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:14:03Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:14:03Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:14:04Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-route [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84296' [e2e-llm-inference-service] uid: 1b210ae9-f1f0-406c-ad72-40f5a8638f73 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-configmap-ref-test [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:03Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:13:56Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] serving.kserve.io/inference-pool-migrated: v1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:serving.kserve.io/inference-pool-migrated: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:14:03Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:14:03Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:14:04Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-route [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84296' [e2e-llm-inference-service] uid: 1b210ae9-f1f0-406c-ad72-40f5a8638f73 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-configmap-ref-test [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:03Z' [e2e-llm-inference-service] message: All references resolved [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:13:56Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 2 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:endpointPickerRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:number: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:matchLabels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPorts: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:14:03Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84277' [e2e-llm-inference-service] uid: 1bd401e3-a7c1-405a-ac37-ce6e2bb23aaf [e2e-llm-inference-service] spec: [e2e-llm-inference-service] appProtocol: http [e2e-llm-inference-service] endpointPickerRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-service [e2e-llm-inference-service] port: [e2e-llm-inference-service] number: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPorts: [e2e-llm-inference-service] - number: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:03Z' [e2e-llm-inference-service] message: Referenced by an HTTPRoute accepted by the parentRef Gateway [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:03Z' [e2e-llm-inference-service] message: Referenced ExtensionRef resolved successfully [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: ResolvedRefs [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: networking.istio.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] kind: AuthPolicy [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:58Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-policies [e2e-llm-inference-service] app.kubernetes.io/managed-by: odh-model-controller [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:rules: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:authentication: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:public: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:anonymous: {} [e2e-llm-inference-service] f:credentials: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:overrides: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fairness: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:response: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:success: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:headers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:x-gateway-inference-fairness-id: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:x-gateway-inference-objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:targetRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:58Z' [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Accepted"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Enforced"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:14:00Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-route-authn [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84242' [e2e-llm-inference-service] uid: 8c5d2221-8a93-4691-818f-e8d5b31fabde [e2e-llm-inference-service] spec: [e2e-llm-inference-service] rules: [e2e-llm-inference-service] authentication: [e2e-llm-inference-service] public: [e2e-llm-inference-service] anonymous: {} [e2e-llm-inference-service] credentials: {} [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] overrides: [e2e-llm-inference-service] fairness: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] objective: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] response: [e2e-llm-inference-service] success: [e2e-llm-inference-service] headers: [e2e-llm-inference-service] x-gateway-inference-fairness-id: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.fairness [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] x-gateway-inference-objective: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.objective [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-route [e2e-llm-inference-service] status: [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:13:59Z' [e2e-llm-inference-service] message: AuthPolicy has been accepted [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:14:00Z' [e2e-llm-inference-service] message: AuthPolicy has been successfully enforced [e2e-llm-inference-service] reason: Enforced [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Enforced [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84121' [e2e-llm-inference-service] uid: a0d177f0-b4cd-4bb6-b320-69e7f36c6a01 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-configmap-ref-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-configmap-ref-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:14:04Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:14:04Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84290' [e2e-llm-inference-service] uid: ff36d9a1-a640-46af-9aa8-25dc663a17e7 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-configmap-ref-test-inference-pool-ip-76890ba8.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-configmap-ref-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84135' [e2e-llm-inference-service] uid: 16b6ffb0-d0fd-4725-a3a3-255bbc42006c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-configmap-ref-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-configmap-ref-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84121' [e2e-llm-inference-service] uid: a0d177f0-b4cd-4bb6-b320-69e7f36c6a01 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-configmap-ref-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-configmap-ref-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:14:04Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:14:04Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84290' [e2e-llm-inference-service] uid: ff36d9a1-a640-46af-9aa8-25dc663a17e7 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-configmap-ref-test-inference-pool-ip-76890ba8.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-configmap-ref-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84135' [e2e-llm-inference-service] uid: 16b6ffb0-d0fd-4725-a3a3-255bbc42006c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-configmap-ref-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-configmap-ref-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84121' [e2e-llm-inference-service] uid: a0d177f0-b4cd-4bb6-b320-69e7f36c6a01 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-configmap-ref-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-configmap-ref-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:14:04Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-shadow-service [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:14:04Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-shadow-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84290' [e2e-llm-inference-service] uid: ff36d9a1-a640-46af-9aa8-25dc663a17e7 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-configmap-ref-test-inference-pool-ip-76890ba8.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-configmap-ref-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84135' [e2e-llm-inference-service] uid: 16b6ffb0-d0fd-4725-a3a3-255bbc42006c [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-configmap-ref-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-configmap-ref-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"d4411d54-bb63-4832-a42f-a86f97d942a7"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:extensionRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:portNumber: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPortNumber: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:13:55Z' [e2e-llm-inference-service] name: scheduler-configmap-ref-test-inference-pool [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-configmap-ref-test [e2e-llm-inference-service] uid: d4411d54-bb63-4832-a42f-a86f97d942a7 [e2e-llm-inference-service] resourceVersion: '84104' [e2e-llm-inference-service] uid: 7dfe4016-ac0d-490a-8da9-2bdd3a8858cd [e2e-llm-inference-service] spec: [e2e-llm-inference-service] extensionRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-configmap-ref-test-epp-service [e2e-llm-inference-service] portNumber: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPortNumber: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parent: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '1970-01-01T00:00:00Z' [e2e-llm-inference-service] message: Waiting for controller [e2e-llm-inference-service] reason: Pending [e2e-llm-inference-service] status: Unknown [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Status [e2e-llm-inference-service] name: default [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:34Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 744bf59fbd [e2e-llm-inference-service] timestamp: '2026-07-02T14:29:12Z' [e2e-llm-inference-service] window: 15.411s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 17419505n [e2e-llm-inference-service] memory: 24756Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:34Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-configmap-ref-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 68f876cc5d [e2e-llm-inference-service] timestamp: '2026-07-02T14:29:13Z' [e2e-llm-inference-service] window: 13.564s [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] usage: [e2e-llm-inference-service] cpu: 11765629n [e2e-llm-inference-service] memory: 30736Ki [e2e-llm-inference-service] apiVersion: metrics.k8s.io/v1beta1 [e2e-llm-inference-service] kind: PodMetrics [e2e-llm-inference-service] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_inference_service] [2026-07-02T14:29:34.838224] end - ❌ 950.965s: ❌ Failed to call model: HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Max retries exceeded with url: /kserve-ci-e2e-test/scheduler-configmap-ref-test/v1/completions (Caused by ReadTimeoutError("HTTPConnectionPool(host='a0b1c0df3ff71427d93c27db9ae1b987-1844924829.us-east-1.elb.amazonaws.com', port=80): Read timed out. (read timeout=60)")) [e2e-llm-inference-service] ---------------------------- Captured log teardown ----------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1708 Deleted ConfigMap scheduler-config-e2e from namespace kserve-ci-e2e-test [e2e-llm-inference-service] _ test_llm_inference_service[router-managed-scheduler-with-replicas-workload-llmd-simulator] _ [e2e-llm-inference-service] [gw0] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] name = 'scheduler-ha-replicas-test', namespace = 'kserve-ci-e2e-test' [e2e-llm-inference-service] version = 'v1alpha1' [e2e-llm-inference-service] [e2e-llm-inference-service] def get_llmisvc( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] name, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] version=constants.KSERVE_V1ALPHA1_VERSION, [e2e-llm-inference-service] ): [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return kserve_client.api_instance.get_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICE, [e2e-llm-inference-service] name, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1043: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test', plural = 'llminferenceservices' [e2e-llm-inference-service] name = 'scheduler-ha-replicas-test', kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] [e2e-llm-inference-service] def get_namespaced_custom_object(self, group, version, namespace, plural, name, **kwargs): # noqa: E501 [e2e-llm-inference-service] """get_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a namespace scoped custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.get_namespaced_custom_object(group, version, namespace, plural, name, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: the custom resource's group (required) [e2e-llm-inference-service] :param str version: the custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: the custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param str name: the custom object's name (required) [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: object [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] kwargs['_return_http_data_only'] = True [e2e-llm-inference-service] > return self.get_namespaced_custom_object_with_http_info(group, version, namespace, plural, name, **kwargs) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:1632: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] group = 'serving.kserve.io', version = 'v1alpha1' [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test', plural = 'llminferenceservices' [e2e-llm-inference-service] name = 'scheduler-ha-replicas-test', kwargs = {'_return_http_data_only': True} [e2e-llm-inference-service] local_var_params = {'_return_http_data_only': True, 'all_params': ['group', 'version', 'namespace', 'plural', 'name', 'async_req', ...], 'auth_settings': ['BearerToken'], 'body_params': None, ...} [e2e-llm-inference-service] all_params = ['group', 'version', 'namespace', 'plural', 'name', 'async_req', ...] [e2e-llm-inference-service] key = '_return_http_data_only', val = True, collection_formats = {} [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'name': 'scheduler-ha-replicas-test', 'namespace': 'kserve-ci-e2e-test', 'plural': 'llminferenceservices', ...} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] def get_namespaced_custom_object_with_http_info(self, group, version, namespace, plural, name, **kwargs): # noqa: E501 [e2e-llm-inference-service] """get_namespaced_custom_object # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] Returns a namespace scoped custom object # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.get_namespaced_custom_object_with_http_info(group, version, namespace, plural, name, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str group: the custom resource's group (required) [e2e-llm-inference-service] :param str version: the custom resource's version (required) [e2e-llm-inference-service] :param str namespace: The custom resource's namespace (required) [e2e-llm-inference-service] :param str plural: the custom resource's plural name. For TPRs this would be lowercase plural kind. (required) [e2e-llm-inference-service] :param str name: the custom object's name (required) [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: tuple(object, status_code(int), headers(HTTPHeaderDict)) [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] local_var_params = locals() [e2e-llm-inference-service] [e2e-llm-inference-service] all_params = [ [e2e-llm-inference-service] 'group', [e2e-llm-inference-service] 'version', [e2e-llm-inference-service] 'namespace', [e2e-llm-inference-service] 'plural', [e2e-llm-inference-service] 'name' [e2e-llm-inference-service] ] [e2e-llm-inference-service] all_params.extend( [e2e-llm-inference-service] [ [e2e-llm-inference-service] 'async_req', [e2e-llm-inference-service] '_return_http_data_only', [e2e-llm-inference-service] '_preload_content', [e2e-llm-inference-service] '_request_timeout' [e2e-llm-inference-service] ] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] for key, val in six.iteritems(local_var_params['kwargs']): [e2e-llm-inference-service] if key not in all_params: [e2e-llm-inference-service] raise ApiTypeError( [e2e-llm-inference-service] "Got an unexpected keyword argument '%s'" [e2e-llm-inference-service] " to method get_namespaced_custom_object" % key [e2e-llm-inference-service] ) [e2e-llm-inference-service] local_var_params[key] = val [e2e-llm-inference-service] del local_var_params['kwargs'] [e2e-llm-inference-service] # verify the required parameter 'group' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('group' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['group'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `group` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'version' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('version' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['version'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `version` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'namespace' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('namespace' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['namespace'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `namespace` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'plural' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('plural' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['plural'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `plural` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] # verify the required parameter 'name' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('name' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['name'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `name` when calling `get_namespaced_custom_object`") # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] collection_formats = {} [e2e-llm-inference-service] [e2e-llm-inference-service] path_params = {} [e2e-llm-inference-service] if 'group' in local_var_params: [e2e-llm-inference-service] path_params['group'] = local_var_params['group'] # noqa: E501 [e2e-llm-inference-service] if 'version' in local_var_params: [e2e-llm-inference-service] path_params['version'] = local_var_params['version'] # noqa: E501 [e2e-llm-inference-service] if 'namespace' in local_var_params: [e2e-llm-inference-service] path_params['namespace'] = local_var_params['namespace'] # noqa: E501 [e2e-llm-inference-service] if 'plural' in local_var_params: [e2e-llm-inference-service] path_params['plural'] = local_var_params['plural'] # noqa: E501 [e2e-llm-inference-service] if 'name' in local_var_params: [e2e-llm-inference-service] path_params['name'] = local_var_params['name'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] [e2e-llm-inference-service] header_params = {} [e2e-llm-inference-service] [e2e-llm-inference-service] form_params = [] [e2e-llm-inference-service] local_var_files = {} [e2e-llm-inference-service] [e2e-llm-inference-service] body_params = None [e2e-llm-inference-service] # HTTP header `Accept` [e2e-llm-inference-service] header_params['Accept'] = self.api_client.select_header_accept( [e2e-llm-inference-service] ['application/json']) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] # Authentication setting [e2e-llm-inference-service] auth_settings = ['BearerToken'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.api_client.call_api( [e2e-llm-inference-service] '/apis/{group}/{version}/namespaces/{namespace}/{plural}/{name}', 'GET', [e2e-llm-inference-service] path_params, [e2e-llm-inference-service] query_params, [e2e-llm-inference-service] header_params, [e2e-llm-inference-service] body=body_params, [e2e-llm-inference-service] post_params=form_params, [e2e-llm-inference-service] files=local_var_files, [e2e-llm-inference-service] response_type='object', # noqa: E501 [e2e-llm-inference-service] auth_settings=auth_settings, [e2e-llm-inference-service] async_req=local_var_params.get('async_req'), [e2e-llm-inference-service] _return_http_data_only=local_var_params.get('_return_http_data_only'), # noqa: E501 [e2e-llm-inference-service] _preload_content=local_var_params.get('_preload_content', True), [e2e-llm-inference-service] _request_timeout=local_var_params.get('_request_timeout'), [e2e-llm-inference-service] collection_formats=collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/custom_objects_api.py:1739: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/{group}/{version}/namespaces/{namespace}/{plural}/{name}' [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] path_params = {'group': 'serving.kserve.io', 'name': 'scheduler-ha-replicas-test', 'namespace': 'kserve-ci-e2e-test', 'plural': 'llminferenceservices', ...} [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], async_req = None, _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def call_api(self, resource_path, method, [e2e-llm-inference-service] path_params=None, query_params=None, header_params=None, [e2e-llm-inference-service] body=None, post_params=None, files=None, [e2e-llm-inference-service] response_type=None, auth_settings=None, async_req=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] """Makes the HTTP request (synchronous) and returns deserialized data. [e2e-llm-inference-service] [e2e-llm-inference-service] To make an async_req request, set the async_req parameter. [e2e-llm-inference-service] [e2e-llm-inference-service] :param resource_path: Path to method endpoint. [e2e-llm-inference-service] :param method: Method to call. [e2e-llm-inference-service] :param path_params: Path parameters in the url. [e2e-llm-inference-service] :param query_params: Query parameters in the url. [e2e-llm-inference-service] :param header_params: Header parameters to be [e2e-llm-inference-service] placed in the request header. [e2e-llm-inference-service] :param body: Request body. [e2e-llm-inference-service] :param post_params dict: Request post form parameters, [e2e-llm-inference-service] for `application/x-www-form-urlencoded`, `multipart/form-data`. [e2e-llm-inference-service] :param auth_settings list: Auth Settings names for the request. [e2e-llm-inference-service] :param response: Response data type. [e2e-llm-inference-service] :param files dict: key -> filename, value -> filepath, [e2e-llm-inference-service] for `multipart/form-data`. [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param collection_formats: dict of collection formats for path, query, [e2e-llm-inference-service] header, and post parameters. [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: [e2e-llm-inference-service] If async_req parameter is True, [e2e-llm-inference-service] the request will be called asynchronously. [e2e-llm-inference-service] The method will return the request thread. [e2e-llm-inference-service] If parameter async_req is False or missing, [e2e-llm-inference-service] then the method will return the response directly. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if not async_req: [e2e-llm-inference-service] > return self.__call_api(resource_path, method, [e2e-llm-inference-service] path_params, query_params, header_params, [e2e-llm-inference-service] body, post_params, files, [e2e-llm-inference-service] response_type, auth_settings, [e2e-llm-inference-service] _return_http_data_only, collection_formats, [e2e-llm-inference-service] _preload_content, _request_timeout, _host) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:348: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceservices/scheduler-ha-replicas-test' [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] path_params = [('group', 'serving.kserve.io'), ('version', 'v1alpha1'), ('namespace', 'kserve-ci-e2e-test'), ('plural', 'llminferenceservices'), ('name', 'scheduler-ha-replicas-test')] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'object' [e2e-llm-inference-service] auth_settings = ['BearerToken'], _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def __call_api( [e2e-llm-inference-service] self, resource_path, method, path_params=None, [e2e-llm-inference-service] query_params=None, header_params=None, body=None, post_params=None, [e2e-llm-inference-service] files=None, response_type=None, auth_settings=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] [e2e-llm-inference-service] config = self.configuration [e2e-llm-inference-service] [e2e-llm-inference-service] # header parameters [e2e-llm-inference-service] header_params = header_params or {} [e2e-llm-inference-service] header_params.update(self.default_headers) [e2e-llm-inference-service] if self.cookie: [e2e-llm-inference-service] header_params['Cookie'] = self.cookie [e2e-llm-inference-service] if header_params: [e2e-llm-inference-service] header_params = self.sanitize_for_serialization(header_params) [e2e-llm-inference-service] header_params = dict(self.parameters_to_tuples(header_params, [e2e-llm-inference-service] collection_formats)) [e2e-llm-inference-service] [e2e-llm-inference-service] # path parameters [e2e-llm-inference-service] if path_params: [e2e-llm-inference-service] path_params = self.sanitize_for_serialization(path_params) [e2e-llm-inference-service] path_params = self.parameters_to_tuples(path_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] for k, v in path_params: [e2e-llm-inference-service] # specified safe chars, encode everything [e2e-llm-inference-service] resource_path = resource_path.replace( [e2e-llm-inference-service] '{%s}' % k, [e2e-llm-inference-service] quote(str(v), safe=config.safe_chars_for_path_param) [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # query parameters [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] query_params = self.sanitize_for_serialization(query_params) [e2e-llm-inference-service] query_params = self.parameters_to_tuples(query_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] # post parameters [e2e-llm-inference-service] if post_params or files: [e2e-llm-inference-service] post_params = post_params if post_params else [] [e2e-llm-inference-service] post_params = self.sanitize_for_serialization(post_params) [e2e-llm-inference-service] post_params = self.parameters_to_tuples(post_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] post_params.extend(self.files_parameters(files)) [e2e-llm-inference-service] [e2e-llm-inference-service] # auth setting [e2e-llm-inference-service] self.update_params_for_auth(header_params, query_params, auth_settings) [e2e-llm-inference-service] [e2e-llm-inference-service] # body [e2e-llm-inference-service] if body: [e2e-llm-inference-service] body = self.sanitize_for_serialization(body) [e2e-llm-inference-service] [e2e-llm-inference-service] # request url [e2e-llm-inference-service] if _host is None: [e2e-llm-inference-service] url = self.configuration.host + resource_path [e2e-llm-inference-service] else: [e2e-llm-inference-service] # use server/host defined in path or operation instead [e2e-llm-inference-service] url = _host + resource_path [e2e-llm-inference-service] [e2e-llm-inference-service] # perform request and return response [e2e-llm-inference-service] > response_data = self.request( [e2e-llm-inference-service] method, url, query_params=query_params, headers=header_params, [e2e-llm-inference-service] post_params=post_params, body=body, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:180: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceservices/scheduler-ha-replicas-test' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] post_params = [], body = None, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] post_params=None, body=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Makes the HTTP request using RESTClient.""" [e2e-llm-inference-service] if method == "GET": [e2e-llm-inference-service] > return self.rest_client.GET(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:373: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceservices/scheduler-ha-replicas-test' [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] query_params = [], _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def GET(self, url, headers=None, query_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] > return self.request("GET", url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] query_params=query_params) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:244: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/serving.kserve.io/v1alpha1/namespaces/kserve-ci-e2e-test/llminferenceservices/scheduler-ha-replicas-test' [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] body=None, post_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Perform requests. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: http request method [e2e-llm-inference-service] :param url: http request url [e2e-llm-inference-service] :param query_params: query parameters in the url [e2e-llm-inference-service] :param headers: http request headers [e2e-llm-inference-service] :param body: request json body, for `application/json` [e2e-llm-inference-service] :param post_params: request post parameters, [e2e-llm-inference-service] `application/x-www-form-urlencoded` [e2e-llm-inference-service] and `multipart/form-data` [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] """ [e2e-llm-inference-service] method = method.upper() [e2e-llm-inference-service] assert method in ['GET', 'HEAD', 'DELETE', 'POST', 'PUT', [e2e-llm-inference-service] 'PATCH', 'OPTIONS'] [e2e-llm-inference-service] [e2e-llm-inference-service] if post_params and body: [e2e-llm-inference-service] raise ApiValueError( [e2e-llm-inference-service] "body parameter cannot be used with post_params parameter." [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] post_params = post_params or {} [e2e-llm-inference-service] headers = headers or {} [e2e-llm-inference-service] [e2e-llm-inference-service] timeout = None [e2e-llm-inference-service] if _request_timeout: [e2e-llm-inference-service] if isinstance(_request_timeout, (int, ) if six.PY3 else (int, long)): # noqa: E501,F821 [e2e-llm-inference-service] timeout = urllib3.Timeout(total=_request_timeout) [e2e-llm-inference-service] elif (isinstance(_request_timeout, tuple) and [e2e-llm-inference-service] len(_request_timeout) == 2): [e2e-llm-inference-service] timeout = urllib3.Timeout( [e2e-llm-inference-service] connect=_request_timeout[0], read=_request_timeout[1]) [e2e-llm-inference-service] [e2e-llm-inference-service] if 'Content-Type' not in headers: [e2e-llm-inference-service] headers['Content-Type'] = 'application/json' [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # For `POST`, `PUT`, `PATCH`, `OPTIONS`, `DELETE` [e2e-llm-inference-service] if method in ['POST', 'PUT', 'PATCH', 'OPTIONS', 'DELETE']: [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] url += '?' + urlencode(query_params) [e2e-llm-inference-service] if (re.search('json', headers['Content-Type'], re.IGNORECASE) or [e2e-llm-inference-service] headers['Content-Type'] == 'application/apply-patch+yaml'): [e2e-llm-inference-service] if headers['Content-Type'] == 'application/json-patch+json': [e2e-llm-inference-service] if not isinstance(body, list): [e2e-llm-inference-service] headers['Content-Type'] = \ [e2e-llm-inference-service] 'application/strategic-merge-patch+json' [e2e-llm-inference-service] request_body = None [e2e-llm-inference-service] if body is not None: [e2e-llm-inference-service] request_body = json.dumps(body) [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'application/x-www-form-urlencoded': # noqa: E501 [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=False, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'multipart/form-data': [e2e-llm-inference-service] # must del headers['Content-Type'], or the correct [e2e-llm-inference-service] # Content-Type which generated by urllib3 will be [e2e-llm-inference-service] # overwritten. [e2e-llm-inference-service] del headers['Content-Type'] [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=True, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] # Pass a `string` parameter directly in the body to support [e2e-llm-inference-service] # other content types than Json when `body` argument is [e2e-llm-inference-service] # provided in serialized form [e2e-llm-inference-service] elif isinstance(body, str) or isinstance(body, bytes): [e2e-llm-inference-service] request_body = body [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Cannot generate the request from given parameters [e2e-llm-inference-service] msg = """Cannot prepare a request message for provided [e2e-llm-inference-service] arguments. Please check that your arguments match [e2e-llm-inference-service] declared content type.""" [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] # For `GET`, `HEAD` [e2e-llm-inference-service] else: [e2e-llm-inference-service] r = self.pool_manager.request(method, url, [e2e-llm-inference-service] fields=query_params, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] except urllib3.exceptions.SSLError as e: [e2e-llm-inference-service] msg = "{0}\n{1}".format(type(e).__name__, str(e)) [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] [e2e-llm-inference-service] if _preload_content: [e2e-llm-inference-service] r = RESTResponse(r) [e2e-llm-inference-service] [e2e-llm-inference-service] # In the python 3, the response.data is bytes. [e2e-llm-inference-service] # we need to decode it to string. [e2e-llm-inference-service] if six.PY3: [e2e-llm-inference-service] r.data = r.data.decode('utf8') [e2e-llm-inference-service] [e2e-llm-inference-service] # log response body [e2e-llm-inference-service] logger.debug("response body: %s", r.data) [e2e-llm-inference-service] [e2e-llm-inference-service] if not 200 <= r.status <= 299: [e2e-llm-inference-service] > raise ApiException(http_resp=r) [e2e-llm-inference-service] E kubernetes.client.exceptions.ApiException: (500) [e2e-llm-inference-service] E Reason: Internal Server Error [e2e-llm-inference-service] E HTTP response headers: HTTPHeaderDict({'Audit-Id': 'e505316d-488e-4927-80b5-377bc18852df', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': 'a8664215-74a2-48d8-b231-e657e37b3200', 'X-Kubernetes-Pf-Prioritylevel-Uid': '0e7454c5-cfcf-4e4f-b47e-e88e460e3769', 'Date': 'Thu, 02 Jul 2026 14:29:45 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] E HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:238: ApiException [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'scheduler-with-replicas', 'workload-llmd-simulator'], prompt='KServe is a', ser... {'name': 'workload-llmd-simulator-schedul-609398f0'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] @pytest.mark.asyncio(loop_scope="session") [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-gateway-ref", [e2e-llm-inference-service] "router-with-managed-route", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[0], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[0]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="custom-route-timeout-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="router-with-refs-test", [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[0], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[0]], [e2e-llm-inference-service] routes=[ROUTER_ROUTES[0], ROUTER_ROUTES[1]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=["router-managed", "workload-pd-cpu", "model-fb-opt-125m"], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-custom-route-timeout-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="custom-route-timeout-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-with-refs-pd", [e2e-llm-inference-service] "scheduler-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="You are an expert in Kubernetes-native machine learning serving platforms, with deep knowledge of the KServe project. " [e2e-llm-inference-service] "Explain the challenges of serving large-scale models, GPU scheduling, and how KServe integrates with capabilities like multi-model serving. " [e2e-llm-inference-service] "Provide a detailed comparison with open source alternatives, focusing on operational trade-offs.", [e2e-llm-inference-service] service_name="router-with-refs-pd-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] expected_gateway=ROUTER_GATEWAYS[1], [e2e-llm-inference-service] before_test=[ [e2e-llm-inference-service] lambda: create_router_resources( [e2e-llm-inference-service] gateways=[ROUTER_GATEWAYS[1]], [e2e-llm-inference-service] routes=[ROUTER_ROUTES[2], ROUTER_ROUTES[3]], [e2e-llm-inference-service] ) [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.custom_gateway, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-dp-ep-gpu", [e2e-llm-inference-service] "workload-dp-ep-prefill-gpu", [e2e-llm-inference-service] "model-deepseek-v2-lite", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="Delve into the multifaceted implications of a fully disaggregated cloud architecture, specifically " [e2e-llm-inference-service] "where the compute plane (P) and the data plane (D) are independently deployed and managed for a " [e2e-llm-inference-service] "geographically distributed, high-throughput, low-latency microservices ecosystem. Beyond the " [e2e-llm-inference-service] "fundamental challenges of network latency and data consistency, elaborate on the advanced " [e2e-llm-inference-service] "considerations and trade-offs inherent in such a setup: 1. Network Architecture and Protocols: " [e2e-llm-inference-service] "How would the network fabric and underlying protocols (e.g., RDMA, custom transport layers) need to " [e2e-llm-inference-service] "evolve to support optimal performance and minimize inter-plane communication overhead, especially for " [e2e-llm-inference-service] "synchronous operations? Discuss the role of network programmability (e.g., SDN, P4) in dynamically " [e2e-llm-inference-service] "optimizing routing and traffic flow between P and D. 2. Advanced Data Consistency and Durability: " [e2e-llm-inference-service] "Explore sophisticated data consistency models (e.g., causal consistency, strong eventual consistency) " [e2e-llm-inference-service] "and their applicability in balancing performance and data integrity across a globally distributed data plane. " [e2e-llm-inference-service] "Detail strategies for ensuring data durability and fault tolerance, including multi-region replication, " [e2e-llm-inference-service] "intelligent partitioning, and recovery mechanisms in the event of partial or full plane failures. " [e2e-llm-inference-service] "3. Dynamic Resource Orchestration and Cost Optimization: Analyze how an orchestration layer would intelligently " [e2e-llm-inference-service] "manage the independent scaling of compute (P) and data (D) resources, considering fluctuating workloads, " [e2e-llm-inference-service] "cost efficiency, and performance targets (e.g., using predictive analytics for resource provisioning). " [e2e-llm-inference-service] "Discuss mechanisms for dynamically reallocating compute nodes to different data partitions based on " [e2e-llm-inference-service] "workload patterns and data locality, potentially involving live migration strategies. " [e2e-llm-inference-service] "4. Security and Compliance in a Distributed Landscape: Address the enhanced security perimeter " [e2e-llm-inference-service] "challenges, including securing communication channels between P and D (encryption in transit, mutual TLS), " [e2e-llm-inference-service] "fine-grained access control to data at rest and in motion, and identity management across disaggregated " [e2e-llm-inference-service] "components. Discuss how such an architecture impacts compliance with regulatory frameworks (e.g., GDPR, HIPAA) " [e2e-llm-inference-service] "concerning data sovereignty, privacy, and auditability. 5. Operational Complexity and Observability: " [e2e-llm-inference-service] "Examine the increased complexity in monitoring, logging, and tracing across highly decoupled compute and " [e2e-llm-inference-service] "data planes. What specialized tooling and practices (e.g., distributed tracing with OpenTelemetry, advanced AIOps) " [e2e-llm-inference-service] "would be essential? How would incident response and troubleshooting differ in this disaggregated environment " [e2e-llm-inference-service] "compared to traditional integrated systems? Consider the challenges of pinpointing root causes across " [e2e-llm-inference-service] "independent failures. 6. Real-world Applicability and Future Trends: Identify specific industries " [e2e-llm-inference-service] "or use cases (e.g., high-frequency trading, IoT edge processing, large language model inference) " [e2e-llm-inference-service] "where the benefits of P/D disaggregation would strongly outweigh its complexities. " [e2e-llm-inference-service] "Conclude by speculating on emerging technologies or paradigms (e.g., serverless compute functions " [e2e-llm-inference-service] "directly interacting with object storage, in-memory disaggregation) that could further drive or " [e2e-llm-inference-service] "transform P/D disaggregation in cloud computing.", [e2e-llm-inference-service] max_tokens=2000, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_gpu, [e2e-llm-inference-service] pytest.mark.cluster_nvidia, [e2e-llm-inference-service] pytest.mark.cluster_nvidia_roce, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-no-scheduler", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.no_scheduler, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="This test simulates DP+EP that can run on CPU, the idea is to test the LWS-based deployment, " [e2e-llm-inference-service] "but without the resources requirements for DP+EP (GPUs and ROCe/IB).", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_multi_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler config tests [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-inline-config-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Chat completions endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] model_name="Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="choices"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-configmap-ref", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-configmap-ref-test", [e2e-llm-inference-service] before_test=[create_scheduler_configmap], [e2e-llm-inference-service] after_test=[delete_scheduler_configmap], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-replicas", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-ha-replicas-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-custom-template", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-custom-template-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Scheduler v0.6 → v0.7 migration tests. [e2e-llm-inference-service] # Deploy v0.6-style configs and verify the controller migrates them [e2e-llm-inference-service] # so the v0.7 scheduler boots successfully. [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-pd-config-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-pd-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-v06-nonzero-threshold-migration", [e2e-llm-inference-service] "workload-llmd-simulator-pd", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="scheduler-v06-threshold-migration-test", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Precise prefix KV cache routing test [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "scheduler-with-precise-prefix-cache-inline-config", [e2e-llm-inference-service] "workload-llmd-simulator-kvcache", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="precise-prefix-cache-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Models endpoint coverage [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=create_response_assertion(with_field="data"), [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/chat/completions [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches("facebook/opt-125m"), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] peers=[ [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-llmd-simulator", [e2e-llm-inference-service] "model-qwen2.5-0.5b", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/chat/completions", [e2e-llm-inference-service] prompt="What is KServe?", [e2e-llm-inference-service] payload_formatter=chat_completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] "Qwen/Qwen2.5-0.5B-Instruct" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/Qwen/Qwen2.5-0.5B-Instruct", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.llmd_simulator, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — LoRA adapter [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/completions", [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] model_name=f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] payload_formatter=completions_payload, [e2e-llm-inference-service] response_assertion=assert_model_field_matches( [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1" [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # Model-based routing via X-Gateway-Model-Name header — /v1/models (base + LoRA) [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m-with-lora-hf", [e2e-llm-inference-service] ], [e2e-llm-inference-service] endpoint="/v1/models", [e2e-llm-inference-service] response_assertion=assert_models_contains( [e2e-llm-inference-service] "facebook/opt-125m", [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] "lora-adapter-1", [e2e-llm-inference-service] f"publishers/{KSERVE_TEST_NAMESPACE}/models/lora-adapter-1", [e2e-llm-inference-service] ), [e2e-llm-inference-service] url_getter=get_model_routing_url, [e2e-llm-inference-service] extra_headers={ [e2e-llm-inference-service] MODEL_ROUTING_HEADER: f"publishers/{KSERVE_TEST_NAMESPACE}/models/facebook/opt-125m", [e2e-llm-inference-service] }, [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.model_routing, [e2e-llm-inference-service] pytest.mark.lora, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] # PVC storage tests -- validate direct PVC volume mount with real vLLM serving [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-pd-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] response_assertion=assert_200_with_choices, [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_single_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-simulated-dp-ep-cpu", [e2e-llm-inference-service] "model-pvc", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] before_test=[ensure_pvc_with_model], [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[ [e2e-llm-inference-service] pytest.mark.cluster_cpu, [e2e-llm-inference-service] pytest.mark.cluster_multi_node, [e2e-llm-inference-service] pytest.mark.pvc_storage, [e2e-llm-inference-service] ], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_llm_inference_service(test_case: TestCase): # noqa: F811 [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] if not test_case.llm_service.metadata.annotations: [e2e-llm-inference-service] test_case.llm_service.metadata.annotations = {} [e2e-llm-inference-service] [e2e-llm-inference-service] test_case.llm_service.metadata.annotations[ [e2e-llm-inference-service] "security.opendatahub.io/enable-auth" [e2e-llm-inference-service] ] = "false" [e2e-llm-inference-service] prefix = test_case.log_prefix [e2e-llm-inference-service] [e2e-llm-inference-service] test_failed = False [e2e-llm-inference-service] try: [e2e-llm-inference-service] print(f"{prefix} Creating LLMInferenceService {service_name}") [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] print(f"{prefix} Waiting for LLMInferenceService {service_name} to be ready") [e2e-llm-inference-service] > wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client, test_case.llm_service, test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:812: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = (, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kin...hedul-1402139b'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-schedul-609398f0'}]}, [e2e-llm-inference-service] 'status': None}, 900) [e2e-llm-inference-service] kwargs = {}, func_name = 'wait_for_llm_isvc_ready' [e2e-llm-inference-service] timestamp_start = '2026-07-02T14:29:35.287727', start_time = 1783002575.2880154 [e2e-llm-inference-service] duration = 9.796285390853882, timestamp_end = '2026-07-02T14:29:45.084304' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] given = {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security....cas-schedul-1402139b'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-schedul-609398f0'}]}, [e2e-llm-inference-service] 'status': None} [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_llm_isvc_ready( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] given: V1alpha1LLMInferenceService, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if "status" not in out: [e2e-llm-inference-service] raise AssertionError("No status found in LLM inference service") [e2e-llm-inference-service] [e2e-llm-inference-service] status = out["status"] [e2e-llm-inference-service] if "conditions" not in status: [e2e-llm-inference-service] raise AssertionError("No conditions found in status") [e2e-llm-inference-service] [e2e-llm-inference-service] expected_true_conditions = {"Ready", "WorkloadsReady", "RouterReady"} [e2e-llm-inference-service] got_true_conditions = set() [e2e-llm-inference-service] [e2e-llm-inference-service] conditions = status["conditions"] [e2e-llm-inference-service] [e2e-llm-inference-service] for condition in conditions: [e2e-llm-inference-service] if condition.get("status") == "True": [e2e-llm-inference-service] got_true_conditions.add(condition.get("type")) [e2e-llm-inference-service] [e2e-llm-inference-service] missing_conditions = expected_true_conditions - got_true_conditions [e2e-llm-inference-service] if missing_conditions: [e2e-llm-inference-service] raise AssertionError( [e2e-llm-inference-service] f"Missing true conditions: {missing_conditions}, expected {expected_true_conditions}, got {conditions}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] return True [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_llm_isvc_ready, timeout=timeout_seconds, interval=1.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1204: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_llm_isvc_ready at 0x7fbae44c8ea0> [e2e-llm-inference-service] timeout = 900, interval = 1.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1215: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_llm_isvc_ready(): [e2e-llm-inference-service] > out = get_llmisvc( [e2e-llm-inference-service] kserve_client, [e2e-llm-inference-service] given.metadata.name, [e2e-llm-inference-service] given.metadata.namespace, [e2e-llm-inference-service] given.api_version.split("/")[1], [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1174: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = [e2e-llm-inference-service] name = 'scheduler-ha-replicas-test', namespace = 'kserve-ci-e2e-test' [e2e-llm-inference-service] version = 'v1alpha1' [e2e-llm-inference-service] [e2e-llm-inference-service] def get_llmisvc( [e2e-llm-inference-service] kserve_client: KServeClient, [e2e-llm-inference-service] name, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] version=constants.KSERVE_V1ALPHA1_VERSION, [e2e-llm-inference-service] ): [e2e-llm-inference-service] try: [e2e-llm-inference-service] return kserve_client.api_instance.get_namespaced_custom_object( [e2e-llm-inference-service] constants.KSERVE_GROUP, [e2e-llm-inference-service] version, [e2e-llm-inference-service] namespace, [e2e-llm-inference-service] KSERVE_PLURAL_LLMINFERENCESERVICE, [e2e-llm-inference-service] name, [e2e-llm-inference-service] ) [e2e-llm-inference-service] except client.rest.ApiException as e: [e2e-llm-inference-service] > raise RuntimeError( [e2e-llm-inference-service] f"❌ Exception when calling CustomObjectsApi->" [e2e-llm-inference-service] f"get_namespaced_custom_object for LLMInferenceService: {e}" [e2e-llm-inference-service] ) from e [e2e-llm-inference-service] E RuntimeError: ❌ Exception when calling CustomObjectsApi->get_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] E Reason: Internal Server Error [e2e-llm-inference-service] E HTTP response headers: HTTPHeaderDict({'Audit-Id': 'e505316d-488e-4927-80b5-377bc18852df', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': 'a8664215-74a2-48d8-b231-e657e37b3200', 'X-Kubernetes-Pf-Prioritylevel-Uid': '0e7454c5-cfcf-4e4f-b47e-e88e460e3769', 'Date': 'Thu, 02 Jul 2026 14:29:45 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] E HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1051: RuntimeError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig router-managed-scheduler-ha-rep-bf59a7b5 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig router-managed-scheduler-ha-rep-bf59a7b5 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig router-managed-scheduler-ha-rep-bf59a7b5 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig scheduler-with-replicas-schedul-1402139b in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig scheduler-with-replicas-schedul-1402139b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig scheduler-with-replicas-schedul-1402139b [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig workload-llmd-simulator-schedul-609398f0 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig workload-llmd-simulator-schedul-609398f0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig workload-llmd-simulator-schedul-609398f0 [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_llm_inference_service] [2026-07-02T14:29:35.220848] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'scheduler-with-replicas', 'workload-llmd-simulator'], prompt='KServe is a', service_name='scheduler-ha-replicas-test', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': None, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'scheduler-ha-replicas-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-scheduler-ha-rep-bf59a7b5'}, [e2e-llm-inference-service] {'name': 'scheduler-with-replicas-schedul-1402139b'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-schedul-609398f0'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-02T14:29:35.233464] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'scheduler-ha-replicas-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-scheduler-ha-rep-bf59a7b5'}, [e2e-llm-inference-service] {'name': 'scheduler-with-replicas-schedul-1402139b'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-schedul-609398f0'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-02T14:29:35.287612] end - ✅ in 0.054s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_llm_isvc_ready] [2026-07-02T14:29:35.287727] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': {'security.opendatahub.io/enable-auth': 'false'}, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'scheduler-ha-replicas-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-scheduler-ha-rep-bf59a7b5'}, [e2e-llm-inference-service] {'name': 'scheduler-with-replicas-schedul-1402139b'}, [e2e-llm-inference-service] {'name': 'workload-llmd-simulator-schedul-609398f0'}]}, [e2e-llm-inference-service] 'status': None}, 900), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: No conditions found in status [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: Missing true conditions: {'WorkloadsReady', 'RouterReady', 'Ready'}, expected {'WorkloadsReady', 'RouterReady', 'Ready'}, got [{'lastTransitionTime': '2026-07-02T14:29:41Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'severity': 'Info', 'status': 'False', 'type': 'HTTPRoutesReady'}, {'lastTransitionTime': '2026-07-02T14:29:41Z', 'message': 'Inference Pool kserve-ci-e2e-test/scheduler-ha-replicas-test-inference-pool exists but no Gateway controller has accepted it yet', 'reason': 'WaitingForGateway', 'severity': 'Info', 'status': 'False', 'type': 'InferencePoolReady'}, {'lastTransitionTime': '2026-07-02T14:29:41Z', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'MainWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:29:41Z', 'severity': 'Info', 'status': 'True', 'type': 'PresetsCombined'}, {'lastTransitionTime': '2026-07-02T14:29:41Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'Ready'}, {'lastTransitionTime': '2026-07-02T14:29:41Z', 'message': 'The following HTTPRoutes are not ready: [kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-route: "False" (reason "InvalidKind", message "referencing unsupported backendRef: group \\"inference.networking.x-k8s.io\\" kind \\"InferencePool\\"")]', 'reason': 'HTTPRoutesNotReady', 'status': 'False', 'type': 'RouterReady'}, {'lastTransitionTime': '2026-07-02T14:29:41Z', 'message': 'Deployment rollout in progress', 'reason': 'Progressing', 'severity': 'Info', 'status': 'False', 'type': 'SchedulerWorkloadReady'}, {'lastTransitionTime': '2026-07-02T14:29:41Z', 'reason': 'Progressing', 'status': 'False', 'type': 'WorkloadsReady'}] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_llm_isvc_ready] [2026-07-02T14:29:45.084304] end - ❌ 9.796s: ❌ Exception when calling CustomObjectsApi->get_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': 'e505316d-488e-4927-80b5-377bc18852df', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': 'a8664215-74a2-48d8-b231-e657e37b3200', 'X-Kubernetes-Pf-Prioritylevel-Uid': '0e7454c5-cfcf-4e4f-b47e-e88e460e3769', 'Date': 'Thu, 02 Jul 2026 14:29:45 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:test_llm_inference_service.py:831 [router-managed-scheduler-with-replicas-workload-llmd-simulator] ❌ ERROR: Failed to call llm inference service scheduler-ha-replicas-test: ❌ Exception when calling CustomObjectsApi->get_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': 'e505316d-488e-4927-80b5-377bc18852df', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': 'a8664215-74a2-48d8-b231-e657e37b3200', 'X-Kubernetes-Pf-Prioritylevel-Uid': '0e7454c5-cfcf-4e4f-b47e-e88e460e3769', 'Date': 'Thu, 02 Jul 2026 14:29:45 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1240 🔍 # Diagnostics for 'scheduler-ha-replicas-test' in 'kserve-ci-e2e-test' [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1241 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1242 # LLMInferenceService scheduler-ha-replicas-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1247 # ❌ failed to dump LLMInferenceService: ❌ Exception when calling CustomObjectsApi->get_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '7e2d603f-cccd-479f-b020-e3cf07357c0f', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': 'a8664215-74a2-48d8-b231-e657e37b3200', 'X-Kubernetes-Pf-Prioritylevel-Uid': '0e7454c5-cfcf-4e4f-b47e-e88e460e3769', 'Date': 'Thu, 02 Jul 2026 14:29:45 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:44 TIME NAMESPACE SOURCE TYPE REASON MESSAGE [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:45 -------------------------------------------------------------------------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-6599bc8fd-8sr69 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.41:8000/health": dial tcp 10.132.0.41:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-6599bc8fd-8sr69 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:56 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:57 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-disabled-test-kserve-router-scheduler-6fcbb55f87-w9mf7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-router-scheduler-6fcbb55f87 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-disabled-test-kserve-6599bc8fd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:52 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy auth-disabled-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "auth-disabled-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-disabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-disabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-disabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-disabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-disabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-disabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-disabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-disabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:56 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-disabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-disabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-67c57d579b-gtwnj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.35/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:05 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 27.649s (27.649s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.35:8000/health": dial tcp 10.132.0.35:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-67c57d579b-gtwnj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.31/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.181s (1.181s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-enabled-test-kserve-router-scheduler-688c79755d-rdsnt [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-router-scheduler-688c79755d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-enabled-test-kserve-67c57d579b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-enabled-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-enabled-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-enabled-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-enabled-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-enabled-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-enabled-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-enabled-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-enabled-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-enabled-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-enabled-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-6cd6c5884-rrbrr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.39:8000/health": dial tcp 10.132.0.39:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-6cd6c5884-rrbrr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler-5b958c76d8sf2gj to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:25 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-router-scheduler-5b958c76d8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set auth-invalid-token-test-kserve-6cd6c5884 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/auth-invalid-token-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/auth-invalid-token-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/auth-invalid-token-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/auth-invalid-token-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:24 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/auth-invalid-token-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/auth-invalid-token-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/auth-invalid-token-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/auth-invalid-token-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:24 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [auth-invalid-token-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:25 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-auth-invalid-token-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:38 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.52:8001/health": dial tcp 10.132.0.52:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-79d6b9cc5-nxhqn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.53/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.53:8000/health": dial tcp 10.132.0.53:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb-g87wg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-prefill-6b6b7bb8bb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-router-scheduler-f7968dxwws to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:34 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-router-scheduler-f7968dfdd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-pd-test-kserve-79d6b9cc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:31 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-pd-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-pd-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:43 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-pd-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:43 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-98989556b-vl89r to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:14:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.45:8000/health": dial tcp 10.132.0.45:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-98989556b-vl89r [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.39/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:01 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:02 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: custom-route-timeout-test-kserve-router-scheduler-579b75452wbbc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-router-scheduler-579b7545d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set custom-route-timeout-test-kserve-98989556b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:57 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy custom-route-timeout-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "custom-route-timeout-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/custom-route-timeout-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/custom-route-timeout-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/custom-route-timeout-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/custom-route-timeout-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/custom-route-timeout-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/custom-route-timeout-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/custom-route-timeout-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:13:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/custom-route-timeout-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:15:01 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [custom-route-timeout-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-custom-route-timeout-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/e2e-pvc-model-download-4s522 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:26 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test job-controller Normal SuccessfulCreate Created pod: e2e-pvc-model-download-4s522 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:36 kserve-ci-e2e-test job-controller Normal Completed Job completed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:20 kserve-ci-e2e-test persistentvolume-controller Normal WaitForFirstConsumer waiting for first consumer to be created before binding [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test persistentvolume-controller Normal ExternalProvisioning Waiting for a volume to be created either by the external provisioner 'ebs.csi.aws.com' or manually by the system administrator. If volume creation is delayed, please verify that the provisioner is running and correctly registered. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:21 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal Provisioning External provisioner is provisioning volume for claim "kserve-ci-e2e-test/e2e-pvc-model-storage" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:23 kserve-ci-e2e-test ebs.csi.aws.com_aws-ebs-csi-driver-controller-544df4d8b9-8r5jw_9b216b93-ea19-4e74-b372-a673518ad683 Normal ProvisioningSucceeded Successfully provisioned volume pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5wjthn to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.094s (1.094s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:23 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:28 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-2f0a622e-kserve-7c9c8cffc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec0c69dceeb48768325d1a53a749e65786-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:22 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-2f0a622e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b9lgcc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.30/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedMount MountVolume.SetUp failed for volume "tls-certs" : secret "gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set gw-section-name-router-with-gat-f1d92d0f-kserve-7bc8dd6c5b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/gw-sec2774c263d49959f50d9eebc552e13bf9-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/gw-section-name-router-with-gat-f1d92d0f-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9499rk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.60/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:46 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.60:8000/health": dial tcp 10.132.0.60:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.60:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-3c960099-kserve-74dfd7fbd9 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:32 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-3c960099-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-3c960099-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-3c960099-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvd10b93c8eba12f87c3a350a5cae2ee0c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:02 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-3c960099] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fbhzphb to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" in 808ms (808ms including waiting). Image size: 44914394 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.48:8001/health": dial tcp 10.132.0.48:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5bcrdvr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.49:8000/health": dial tcp 10.132.0.49:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill-5b86594bc5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-50bc673d-kserve-6cb68959fb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:44 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:14 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv44d181485fad85e662eb092f3749502f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-50bc673d-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:52 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-50bc673d] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test00d7278d8a22c4e39146a6b0eb840f45-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:07:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.44:8000/health": dial tcp 10.132.0.44:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d-qqjcw [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-87882a8e-kserve-bb656dd4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisva690bbc929faec8bc98c767f16c003c1-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:06 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-87882a8e-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:07:54 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-87882a8e] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test21fe6730fe484f3a92b1a16afe1bac8f-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" in 44.193s (44.193s including waiting). Image size: 3531177328 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0-1 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.62/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:42 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:40:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.62:8000/health": dial tcp 10.132.0.62:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 7f6c8cfc6 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 in StatefulSet llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:40:17 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test386bd5808c5c450f11fd8632e1f821fb-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-dc21cb14-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:58 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-dc21cb14] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:44 kserve-ci-e2e-test LLMInferenceServiceController Warning Error Reconciliation failed: failed to reconcile networking: failed to reconcile scheduler: failed to build expected scheduler deployment: failed to attach model artifacts to scheduler deployment: Get "https://172.31.0.1:443/api/v1/namespaces/kserve-ci-e2e-test/secrets/llmisvc-model-fb-opt-125m-route-dc21cb14-epp-sa-dockercfg-l5xfh": context canceled [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:44 kserve-ci-e2e-test LLMInferenceServiceController Warning UpdateFailed Failed to update status for LLMInferenceService "llmisvc-model-fb-opt-125m-route-dc21cb14": context canceled [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d49gs9g to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:20 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.45:8000/health": context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-66874c76d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:27 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:31 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:32 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-route-e95b1dc1-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv122f03714c5bdf915a2917fdf1262b98-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:59 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-route-e95b1dc1] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.42:8000/health": dial tcp 10.132.0.42:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5-grs4g [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler-67f9b884d4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-schet7zjm to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:05 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-7ca60146-kserve-c44fc98d5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:01 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:19 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv3e414c2ba058a022dfd694dbcbac5b51-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:01:20 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-7ca60146-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-7ca60146] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test5216bfd716f919dc046bc693ceb22e41-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:44 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.43:8000/health": dial tcp 10.132.0.43:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb-mw6jk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:35 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-sche2jmxn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler-5444b4dbf from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-586f8c4ddb from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:57 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv77ff2528d3e9b4972cd9335229fce9f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:58 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-fb-opt-125m-with-ba4d693a-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:06:55 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-fb-opt-125m-with-ba4d693a] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test05addb65ba05195619f26ef266e8fc04-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.55/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.54/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:25 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.54:8000/health": dial tcp 10.132.0.54:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal CreatingRevision Creating revision with key 686d468674 for a newly created LeaderWorkerSet [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created leader statefulset llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Replicas are progressing, with 0 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test statefulset-controller Normal SuccessfulCreate create Pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 in StatefulSet llmisvc-model-pvc-router-manage-2577e794-kserve-mn successful [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test leaderworkerset Normal GroupsProgressing Created worker statefulset for leader pod llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:35 kserve-ci-e2e-test leaderworkerset Normal AllGroupsReady All replicas are ready, with 1 groups ready of total 1 groups [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-2577e794-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-2577e794-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn-scc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.LeaderWorkerSet kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-mn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv1dc4269d1ada5f2d28562215d180c57f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:27:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-2577e794-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:53 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-2577e794] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:56 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test37591b20e96e9663d45a730d03070f1e-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:15 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.47:8000/health": dial tcp 10.132.0.47:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:51 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-59b9d263-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-59b9d263-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv8bf079eb6eda4debfb4ef5bb7817824c-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:36 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-59b9d263-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:25 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-59b9d263] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testc9569cf4801efc0ed27b2f25ffaee875-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.50:8001/health": dial tcp 10.132.0.50:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Normal SuccessfulAttachVolume AttachVolume.Attach succeeded for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:25:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.51:8000/health": dial tcp 10.132.0.51:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-79497db4cc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:56 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-pvc-router-manage-e8706282-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-pvc-router-manage-e8706282-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb19f98874e050eec8ca94d49676113f0-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-pvc-router-manage-e8706282-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:09 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-pvc-router-manage-e8706282] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testf5d060a5eb39a04e074b78907a1556a6-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff448t5z to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.51/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-797dfc7ff4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test80f99418f0f3eb4c80f02e6144a6d4f2-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv431d7e30f13554296364403a5b6365f5-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:36 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4c1ba212] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98ffcrwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:30 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.34:8000/health": dial tcp 10.133.0.34:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-5c54ddb98f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:36 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv08544b88a8d963ffd553cc1f3ed82d16-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-4f8c0978-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:14 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-4f8c0978] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test76d7f532acb694e4a7bcef75d32cd8a1-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd8mtrg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.32/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-5d8ffd58dd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:50 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:01 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvb11a9c9832b99b016bc8f8e0ea095712-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-model-qwen2-5-0-5b-rout-a50492e9-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:48 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-model-qwen2-5-0-5b-rout-a50492e9] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testb7025ca4d8a6f8f5b2fd08b5581d2678-llmisvc-mode [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.43/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56-hxmqv [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-4b931143-kserve-bd545d56 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:40 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-4b931143-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-4b931143-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:48 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisvca2d2d7d499abb359505529ebe02c136-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-4b931143-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:21 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-4b931143] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test8ac8e3d2264ccb939eb021b0b835847c-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb478lth to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:25 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.134.0.42:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-5b1e8f15-kserve-64df7bddb4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-5b1e8f15-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-5b1e8f15-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:49 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisve55ae740357a3a31a27cdb8b66ffe20f-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-5b1e8f15-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:22 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-5b1e8f15] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test7f54e84970003a6e7372bdbcb574f7ed-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879srkgh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc-router-managed-test-llm-e45d1f79-kserve-7fdbbd4879 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy llmisvc-router-managed-test-llm-e45d1f79-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route: HTTPRoute.gateway.networking.k8s.io "llmisvc-router-managed-test-llm-e45d1f79-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/llmisv5c7e67b6c51568d1d6d13829a9337f2a-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/llmisvc-router-managed-test-llm-e45d1f79-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:20 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [llmisvc-router-managed-test-llm-e45d1f79] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-testef4d2875be14b30dc1561ed84d0d4bde-llmisvc-rout [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-scheduler-6fcb489785 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc08544b88a8d963ffd553cc1f3ed82d16-kserve-router-schelgz65 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.566s (1.566s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:59:42 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:15 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-scheduler-7dbcb75dbc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc1dc4269d1ada5f2d28562215d180c57f-kserve-router-schej697x to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.48/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:48 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:49 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:50 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0, llmisvc-model-pvc-router-manage-2577e794-kserve-mn-0-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:30:57 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-scheduler-5688c7c666 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc32fd5468e8138ec3cd54cf7b9c18cfb0-kserve-router-schempqgn to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:37 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:02 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-scheduler-79866dc455 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc431d7e30f13554296364403a5b6365f5-kserve-router-schetml5l to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:03 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedToRetrieveImagePullSecret Unable to retrieve some image pull secrets (llmisvc-model-qwen2-5-0-5b-rout-4c1ba212-epp-sa-dockercfg-pg2h6); attempting to pull the image may not succeed. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:58:04 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc44d181485fad85e662eb092f3749502f-kserve-router-scheduler-69bf65579d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc44d181485fad85e662eb092f3749502f-kserve-router-schepbszb to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.42/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.36/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:48 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-sche2j98f [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:47 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc5c7e67b6c51568d1d6d13829a9337f2a-kserve-router-scheduler-5dd88bfbb7 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:11 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-59b9d263-kserve-585587bc9dghndz [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:30 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-sche8wvg6 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:09 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvc8bf079eb6eda4debfb4ef5bb7817824c-kserve-router-scheduler-5f555d4d85 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:55 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:12:26 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-sche667dg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:05:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvca690bbc929faec8bc98c767f16c003c1-kserve-router-scheduler-54ccf64f4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-scheduler-6d86bd4d9d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb11a9c9832b99b016bc8f8e0ea095712-kserve-router-schernxht to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.38/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:56:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:57:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:01 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-scheduler-67b4bb9646 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcb19f98874e050eec8ca94d49676113f0-kserve-router-schep854q to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:02 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:22:04 kserve-ci-e2e-test attachdetach-controller Warning FailedAttachVolume Multi-Attach error for volume "pvc-c7f50f50-4c3f-4681-9e32-a00c68b5f472" Volume is already used by pod(s) llmisvc-model-pvc-router-manage-e8706282-kserve-7db79954c-9trd9, llmisvc-model-pvc-router-manage-e8706282-kserve-prefill-795d4gk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:26:17 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-scheduler-68cc9685d6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvcca2d2d7d499abb359505529ebe02c136-kserve-router-schembsjj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.40/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:00:35 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.37/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:58:09 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-sche8kwzk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:55:50 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set llmisvce55ae740357a3a31a27cdb8b66ffe20f-kserve-router-scheduler-749449dbc8 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-ccbb969bd-65ln7 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.63/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:12 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:56:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.63:8000/health": dial tcp 10.132.0.63:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-multiple-adapters-test-kserve-ccbb969bd-65ln7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-multiple-adapters-test-kserve-ccbb969bd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:08 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-multiple-adapters-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-multiple-adapters-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-multiple-adapters-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:11 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:55:17 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-multiple-adapters-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:56:42 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-multiple-adapters-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.61/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:06 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:21 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:52 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.61:8000/health": dial tcp 10.132.0.61:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: lora-single-adapter-test-kserve-85cd6c88dc-6zqb8 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set lora-single-adapter-test-kserve-85cd6c88dc from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy lora-single-adapter-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "lora-single-adapter-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/lora-single-adapter-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.HTTPRoute kserve-ci-e2e-test/lora-single-adapter-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:00 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.DestinationRule kserve-ci-e2e-test/lora-single-adapter-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:38:10 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/lora-single-adapter-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:39:31 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [lora-single-adapter-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.29/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.483s (1.483s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:33 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.133.0.29:8000/health": net/http: request canceled while waiting for connection (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.33/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 3.651s (3.651s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:11 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" in 1.88s (1.88s including waiting). Image size: 98346788 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-rtxvp [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-7c94f9d7c6-p6h6p [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.34/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" in 2.962s (2.962s including waiting). Image size: 301878719 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:07 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:08 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" in 1.723s (1.723s including waiting). Image size: 75073927 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulling Pulling image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Successfully pulled image "ghcr.io/llm-d/llm-d-uds-tokenizer:vllm-v0.19.1" in 30.447s (30.447s including waiting). Image size: 2989890188 bytes. [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:40 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:41 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Liveness probe failed: timeout: failed to connect service "10.132.0.34:9003" within 1s: context deadline exceeded [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container tokenizer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:17 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: precise-prefix-cache-test-kserve-router-scheduler-77d88bdc8mp29 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-router-scheduler-77d88bdcf4 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set precise-prefix-cache-test-kserve-7c94f9d7c6 from 0 to 2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy precise-prefix-cache-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "precise-prefix-cache-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/precise-prefix-cache-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/precise-prefix-cache-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/precise-prefix-cache-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/precise-prefix-cache-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/precise-prefix-cache-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/precise-prefix-cache-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/precise-prefix-cache-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:27 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/precise-prefix-cache-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:15 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [precise-prefix-cache-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:54:16 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-precise-prefix-cache-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/prestop-hook-test-kserve-66474c4c54-j59xs to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.67/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:04 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.67:8000/health": dial tcp 10.132.0.67:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: prestop-hook-test-kserve-66474c4c54-j59xs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/prestop-hook-test-kserve-router-scheduler-7c7698955f-mvsc8 to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:00 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.68/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:00 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:01 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: prestop-hook-test-kserve-router-scheduler-7c7698955f-mvsc8 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set prestop-hook-test-kserve-router-scheduler-7c7698955f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set prestop-hook-test-kserve-66474c4c54 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/prestop-hook-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/prestop-hook-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/prestop-hook-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/prestop-hook-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-prestop-hook-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/prestop-hook-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/prestop-hook-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/prestop-hook-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:59 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/prestop-hook-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/prestop-hook-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:07 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/prestop-hook-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:08 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/prestop-hook-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-1-openshift-default-799f46c59b-fnbwc to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.28/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:31 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Readiness probe failed: Get "http://10.133.0.28:15021/healthz/ready": context deadline exceeded (Client.Timeout exceeded while awaiting headers) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:02:52 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning BackOff Back-off restarting failed container istio-proxy in pod router-gateway-1-openshift-default-799f46c59b-fnbwc_kserve-ci-e2e-test(f244ca92-1cf1-429e-886d-e0d585d0175f) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:03:09 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.28:15021/healthz/ready": dial tcp 10.133.0.28:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-1-openshift-default-799f46c59b-fnbwc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:58 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-1-openshift-default-799f46c59b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:52:59 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 12:53:03 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:37 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-gateway-2-openshift-default-54c789bdc6-w57g9 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.44/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "registry.redhat.io/openshift-service-mesh/istio-proxyv2-rhel9@sha256:7d15cebf9b62f3f235c0eab5158ac8ff2fda86a1d193490dc94c301402c99da8" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container istio-proxy [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:19 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning Unhealthy Startup probe failed: Get "http://10.133.0.44:15021/healthz/ready": dial tcp 10.133.0.44:15021: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-gateway-2-openshift-default-54c789bdc6-w57g9 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-gateway-2-openshift-default-54c789bdc6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test service-controller Normal EnsuringLoadBalancer Ensuring load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:17 kserve-ci-e2e-test service-controller Normal EnsuredLoadBalancer Ensured load balancer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:12 kserve-ci-e2e-test gateway_labeler_controller Normal AddedLabel Added label istio.io/rev=openshift-gateway to gateway router-gateway-2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-disagg-sidecar:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:32 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.56:8001/health": dial tcp 10.132.0.56:8001: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container llm-d-routing-sidecar [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-694c6ddf4c-5rnrl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.57/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.57:8000/health": dial tcp 10.132.0.57:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-prefill-778968b98b-z5qsg [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-prefill-778968b98b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-pd-test-kserve-router-scheduler-5c684fd54fjjls to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.45/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:03 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:27 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-router-scheduler-5c684fd54d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-pd-test-kserve-694c6ddf4c from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-pd-test-kserve-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-pd-test-kserve-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-pd-test-kserve-prefill [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-pd-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-pd-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:28 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-pd-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:33 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-pd-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:47 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-pd-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:02 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-pd-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-587fcd8566-2pqvl to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:50 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:20:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:29 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.46:8000/health": dial tcp 10.132.0.46:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-587fcd8566-2pqvl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.41/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:47 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: router-with-refs-test-kserve-router-scheduler-74545d8489-mw8dl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-router-scheduler-74545d8489 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set router-with-refs-test-kserve-587fcd8566 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:44 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/router-with-refs-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/router-with-refs-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/router-with-refs-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/router-with-refs-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:45 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/router-with-refs-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:18:46 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/router-with-refs-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:19:23 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/router-with-refs-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:05 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [router-with-refs-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:21:13 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-router-with-refs-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.56/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-configmap-ref-test-kserve-744bf59fbd-q9fjk [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-router-scheduler-68f87fnzqr to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.66/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-configmap-ref-test-kserve-router-scheduler-68f876cc5d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-configmap-ref-test-kserve-744bf59fbd from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy scheduler-configmap-ref-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "scheduler-configmap-ref-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-configmap-ref-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:55 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-configmap-ref-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:03 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/scheduler-configmap-ref-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:14:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [scheduler-configmap-ref-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-5c5cfcf776-sp8jl to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.52/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-ha-replicas-test-kserve-5c5cfcf776-sp8jl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85d7nl6t to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:39 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.69/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:39 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85d7nl6t [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dhvggl [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dhvggl to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:39 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.57/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:39 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:39 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:39 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dbc4 from 0 to 2 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-ha-replicas-test-kserve-5c5cfcf776 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:37 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy scheduler-ha-replicas-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "scheduler-ha-replicas-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-ha-replicas-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/scheduler-ha-replicas-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/scheduler-ha-replicas-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:38 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-ha-replicas-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:29:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-7c6c58cc96-l62qh to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.49/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-inference-sim:v0.8.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-7c6c58cc96-l62qh [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler-d9ddfhjjld to ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.134.0.50/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:40 kserve-ci-e2e-test kubelet/ip-10-0-138-159.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-router-scheduler-d9ddf6c4d from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set scheduler-inline-config-test-kserve-7c6c58cc96 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:45 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy scheduler-inline-config-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "scheduler-inline-config-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:53 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/scheduler-inline-config-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/scheduler-inline-config-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/scheduler-inline-config-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/scheduler-inline-config-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:53:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/scheduler-inline-config-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/scheduler-inline-config-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:04 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/scheduler-inline-config-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:05 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/scheduler-inline-config-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:54:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [scheduler-inline-config-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:57:39 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-scheduler-inline-config-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-6hzlt to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.59/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:18 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:24 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:51 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.59:8000/health": dial tcp 10.132.0.59:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-5d797b476b-f2dmd to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.58/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:47 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:43 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:10 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.58:8000/health": dial tcp 10.132.0.58:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-f2dmd [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-5d797b476b-6hzlt [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.46/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:43 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:44 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d to ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.133.0.47/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:15 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:36 kserve-ci-e2e-test kubelet/ip-10-0-141-170.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-5s4p7 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: stop-feature-test-kserve-router-scheduler-69f4578b7f-jpj2d [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-router-scheduler-69f4578b7f from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:14 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set stop-feature-test-kserve-5d797b476b from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:35:07 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy stop-feature-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "stop-feature-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-stop-feature-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/stop-feature-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/stop-feature-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:42 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/stop-feature-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:31:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/stop-feature-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:37:34 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [stop-feature-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Secret kserve-ci-e2e-test/stop-feature-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ServiceAccount kserve-ci-e2e-test/stop-feature-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Deployment kserve-ci-e2e-test/stop-feature-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.Service kserve-ci-e2e-test/stop-feature-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:54 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.InferencePool kserve-ci-e2e-test/stop-feature-test-inference-pool [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 13:34:56 kserve-ci-e2e-test LLMInferenceServiceController Warning LLMInferenceServiceNotReady LLMInferenceService [stop-feature-test] is no longer Ready because of: GatewaysReady, HTTPRoutesReady, InferencePoolReady, MainWorkloadReady, PrefillWorkerWorkloadReady, PrefillWorkloadReady, RouterReady, SchedulerWorkloadReady, WorkerWorkloadReady, WorkloadsReady [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/tls-verification-test-kserve-6d6946dcd6-bdddw to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.64/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "public.ecr.aws/q9t5s3a7/vllm-cpu-release-repo:v0.19.0" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:31 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:16 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Startup probe failed: Get "https://10.132.0.64:8000/health": dial tcp 10.132.0.64:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:49 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning Unhealthy Readiness probe failed: Get "https://10.132.0.64:8000/health": dial tcp 10.132.0.64:8000: connect: connection refused [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: tls-verification-test-kserve-6d6946dcd6-bdddw [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 None kserve-ci-e2e-test Normal Scheduled Successfully assigned kserve-ci-e2e-test/tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj to ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test multus Normal AddedInterface Add eth0 [10.132.0.65/23] from ovn-kubernetes [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "quay.io/opendatahub/kserve-storage-initializer@sha256:ad817336055c6da6bf66cacbd06b857caf8adb146e60422365f2b38a3444569b" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:27 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container storage-initializer [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Pulled Container image "ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2" already present on machine [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Created Created container: main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:28 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Started Started container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Normal Killing Stopping container main [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:34 kserve-ci-e2e-test kubelet/ip-10-0-134-1.ec2.internal Warning FailedPreStopHook PreStopHook failed [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test replicaset-controller Normal SuccessfulCreate Created pod: tls-verification-test-kserve-router-scheduler-7f955879f5-96mlj [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set tls-verification-test-kserve-router-scheduler-7f955879f5 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test deployment-controller Normal ScalingReplicaSet Scaled up replica set tls-verification-test-kserve-6d6946dcd6 from 0 to 1 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:24 kserve-ci-e2e-test OpenDataHubModelController Warning ReconcileError Failed to reconcile LLMInferenceService: 1 error occurred: * failed to get HTTPRoute for AuthPolicy tls-verification-test-kserve-route-authn: failed to get HTTPRoute kserve-ci-e2e-test/tls-verification-test-kserve-route: HTTPRoute.gateway.networking.k8s.io "tls-verification-test-kserve-route" not found [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Secret kserve-ci-e2e-test/tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/tls-verification-test-kserve [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/tls-verification-test-kserve-workload-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ServiceAccount kserve-ci-e2e-test/tls-verification-test-epp-sa [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.ClusterRoleBinding /kserve-ci-e2e-test-tls-verification-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Role kserve-ci-e2e-test/tls-verification-test-epp-role [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.RoleBinding kserve-ci-e2e-test/tls-verification-test-epp-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Deployment kserve-ci-e2e-test/tls-verification-test-kserve-router-scheduler [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:26 kserve-ci-e2e-test LLMInferenceServiceController Normal Created Created v1.Service kserve-ci-e2e-test/tls-verification-test-epp-service [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Created (combined from similar events): Created v1.DestinationRule kserve-ci-e2e-test/tls-verification-test-kserve-shadow-svc [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:51 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.Secret kserve-ci-e2e-test/tls-verification-test-kserve-self-signed-certs [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:11:52 kserve-ci-e2e-test LLMInferenceServiceController Normal Updated Updated v1.HTTPRoute kserve-ci-e2e-test/tls-verification-test-kserve-route [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:13:27 kserve-ci-e2e-test LLMInferenceServiceController Normal LLMInferenceServiceReady LLMInferenceService [tls-verification-test] is Ready [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:56 2026-07-02 14:28:34 kserve-ci-e2e-test LLMInferenceServiceController Normal Deleted Deleted v1.ClusterRoleBinding /kserve-ci-e2e-test-tls-verification-test-epp-auth-rb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod scheduler-ha-replicas-test-kserve-5c5cfcf776-sp8jl (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:204 # -- logs (current): unavailable ((500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': 'ecbe8552-2600-494e-9f98-634fc1fba789', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'Date': 'Thu, 02 Jul 2026 14:29:46 GMT', 'Content-Length': '235'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"Get \"https://10.0.141.170:10250/containerLogs/kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-5c5cfcf776-sp8jl/main?tailLines=200\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] ) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85d7nl6t (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:204 # -- logs (current): unavailable ((500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '6e30f288-b2d9-499f-9ae0-d7a402c9f6e6', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'Date': 'Thu, 02 Jul 2026 14:29:46 GMT', 'Content-Length': '246'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"Get \"https://10.0.134.1:10250/containerLogs/kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85d7nl6t/main?tailLines=200\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] ) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:148 ### Pod scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dhvggl (phase=Running) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:188 #### container 'main' (restarts=0) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:diagnostic.py:204 # -- logs (current): unavailable ((500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': '2b664244-3b21-4822-b172-3a54be6eca05', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'Date': 'Thu, 02 Jul 2026 14:29:46 GMT', 'Content-Length': '248'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"Get \"https://10.0.138.159:10250/containerLogs/kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dhvggl/main?tailLines=200\": EOF","code":500} [e2e-llm-inference-service] [e2e-llm-inference-service] ) [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-service [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 39d2e4a8-3c84-4851-b95b-eb9cf898f7ea [e2e-llm-inference-service] resourceVersion: '92155' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - notReadyAddresses: [e2e-llm-inference-service] - ip: 10.132.0.69 [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85d7nl6t [e2e-llm-inference-service] uid: 463c658e-bddd-4fd5-a086-b17fc708ddce [e2e-llm-inference-service] - ip: 10.134.0.57 [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dhvggl [e2e-llm-inference-service] uid: 204bdfda-9eed-4bad-880b-7f21471eae74 [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 652182ab-9cff-4984-856b-40ac7ce1d5aa [e2e-llm-inference-service] resourceVersion: '92145' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpoints.kubernetes.io/managed-by: endpoint-controller [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:subsets: {} [e2e-llm-inference-service] subsets: [e2e-llm-inference-service] - notReadyAddresses: [e2e-llm-inference-service] - ip: 10.133.0.52 [e2e-llm-inference-service] nodeName: ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-5c5cfcf776-sp8jl [e2e-llm-inference-service] uid: 0e8fac6c-4313-4304-b55c-30ffeecc7fff [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Endpoints [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-5c5cfcf776-sp8jl [e2e-llm-inference-service] generateName: scheduler-ha-replicas-test-kserve-5c5cfcf776- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 0e8fac6c-4313-4304-b55c-30ffeecc7fff [e2e-llm-inference-service] resourceVersion: '92142' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 5c5cfcf776 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.133.0.52/23"],"mac_address":"0a:58:0a:85:00:34","gateway_ips":["10.133.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.133.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.133.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.133.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.133.0.1"}],"ip_address":"10.133.0.52/23","gateway_ip":"10.133.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.133.0.52\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:85:00:34\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-5c5cfcf776 [e2e-llm-inference-service] uid: 0e32e010-9a87-43e0-89ce-35a224a99541 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-141-170 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"0e32e010-9a87-43e0-89ce-35a224a99541"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.133.0.52"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-ha-replicas-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: kube-api-access-c8cx4 [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/llm-d-inference-sim [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --model [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --mode [e2e-llm-inference-service] - random [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: kube-api-access-c8cx4 [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: default [e2e-llm-inference-service] serviceAccount: default [e2e-llm-inference-service] nodeName: ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] fsGroup: 1000690000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] hostIP: 10.0.141.170 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.141.170 [e2e-llm-inference-service] podIP: 10.133.0.52 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.133.0.52 [e2e-llm-inference-service] startTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: false [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] imageID: ghcr.io/llm-d/llm-d-inference-sim@sha256:bab162bd25e2ed8b15022387cdb223023aeb33be49476af9f0115c0398fb8ff5 [e2e-llm-inference-service] containerID: cri-o://110b2809bba5850c81f5db9102d712182917d3eae1c48d6076336ce4b54d0461 [e2e-llm-inference-service] started: false [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: kube-api-access-c8cx4 [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85d7nl6t [e2e-llm-inference-service] generateName: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dbc4- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 463c658e-bddd-4fd5-a086-b17fc708ddce [e2e-llm-inference-service] resourceVersion: '92153' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 6f5f85dbc4 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.132.0.69/23"],"mac_address":"0a:58:0a:84:00:45","gateway_ips":["10.132.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.132.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.132.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.132.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.132.0.1"}],"ip_address":"10.132.0.69/23","gateway_ip":"10.132.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.132.0.69\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:84:00:45\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dbc4 [e2e-llm-inference-service] uid: b92c1a65-aee7-4258-8de7-ecd6bc2982cc [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-134-1 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"b92c1a65-aee7-4258-8de7-ecd6bc2982cc"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.132.0.69"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-ha-replicas-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kube-api-access-xt52w [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --ha-enable-leader-election [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type: prefix-cache-scorer\n\ [e2e-llm-inference-service] - type: max-score-picker\nschedulingProfiles:\n- name: default\n plugins:\n\ [e2e-llm-inference-service] \ - pluginRef: queue-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n\ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-xt52w [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] serviceAccount: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] fsGroup: 1000690000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: scheduler-ha-replicas-test-epp-sa-dockercfg-mp5wd [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] hostIP: 10.0.134.1 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.134.1 [e2e-llm-inference-service] podIP: 10.132.0.69 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.132.0.69 [e2e-llm-inference-service] startTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: false [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] imageID: ghcr.io/llm-d/llm-d-router-endpoint-picker@sha256:06b6c75d77afd0e07053402752a9736c2dfbc12a306d0d37d963aac4c1d4e6a6 [e2e-llm-inference-service] containerID: cri-o://183b5242879c45427b1f2068c105744b2f63a981378f9ed7f7de01a310db8425 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-xt52w [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dhvggl [e2e-llm-inference-service] generateName: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dbc4- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 204bdfda-9eed-4bad-880b-7f21471eae74 [e2e-llm-inference-service] resourceVersion: '92141' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 6f5f85dbc4 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] k8s.ovn.org/pod-networks: '{"default":{"ip_addresses":["10.134.0.57/23"],"mac_address":"0a:58:0a:86:00:39","gateway_ips":["10.134.0.1"],"routes":[{"dest":"10.132.0.0/14","nextHop":"10.134.0.1"},{"dest":"172.31.0.0/16","nextHop":"10.134.0.1"},{"dest":"169.254.0.5/32","nextHop":"10.134.0.1"},{"dest":"100.64.0.0/16","nextHop":"10.134.0.1"}],"ip_address":"10.134.0.57/23","gateway_ip":"10.134.0.1","role":"primary"}}' [e2e-llm-inference-service] k8s.v1.cni.cncf.io/network-status: "[{\n \"name\": \"ovn-kubernetes\",\n \ [e2e-llm-inference-service] \ \"interface\": \"eth0\",\n \"ips\": [\n \"10.134.0.57\"\n ],\n\ [e2e-llm-inference-service] \ \"mac\": \"0a:58:0a:86:00:39\",\n \"default\": true,\n \"dns\": {}\n\ [e2e-llm-inference-service] }]" [e2e-llm-inference-service] openshift.io/scc: restricted-v2 [e2e-llm-inference-service] seccomp.security.alpha.kubernetes.io/pod: runtime/default [e2e-llm-inference-service] security.openshift.io/validated-scc-subject-type: user [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dbc4 [e2e-llm-inference-service] uid: b92c1a65-aee7-4258-8de7-ecd6bc2982cc [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: ip-10-0-138-159 [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.ovn.org/pod-networks: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"b92c1a65-aee7-4258-8de7-ecd6bc2982cc"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:enableServiceLinks: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kubelet [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] k:{"type":"ContainersReady"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Initialized"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodReadyToStartContainers"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"PodScheduled"}: [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] k:{"type":"Ready"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastProbeTime: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:containerStatuses: {} [e2e-llm-inference-service] f:hostIP: {} [e2e-llm-inference-service] f:hostIPs: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:phase: {} [e2e-llm-inference-service] f:podIP: {} [e2e-llm-inference-service] f:podIPs: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"ip":"10.134.0.57"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:ip: {} [e2e-llm-inference-service] f:startTime: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: multus-daemon [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:k8s.v1.cni.cncf.io/network-status: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-ha-replicas-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: kube-api-access-k87xh [e2e-llm-inference-service] projected: [e2e-llm-inference-service] sources: [e2e-llm-inference-service] - serviceAccountToken: [e2e-llm-inference-service] expirationSeconds: 3607 [e2e-llm-inference-service] path: token [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: kube-root-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: ca.crt [e2e-llm-inference-service] path: ca.crt [e2e-llm-inference-service] - downwardAPI: [e2e-llm-inference-service] items: [e2e-llm-inference-service] - path: namespace [e2e-llm-inference-service] fieldRef: [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] fieldPath: metadata.namespace [e2e-llm-inference-service] - configMap: [e2e-llm-inference-service] name: openshift-service-ca.crt [e2e-llm-inference-service] items: [e2e-llm-inference-service] - key: service-ca.crt [e2e-llm-inference-service] path: service-ca.crt [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --ha-enable-leader-election [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type: prefix-cache-scorer\n\ [e2e-llm-inference-service] - type: max-score-picker\nschedulingProfiles:\n- name: default\n plugins:\n\ [e2e-llm-inference-service] \ - pluginRef: queue-scorer\n weight: 2\n - pluginRef: prefix-cache-scorer\n\ [e2e-llm-inference-service] \ weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-k87xh [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsUser: 1000690000 [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] serviceAccount: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] seLinuxOptions: [e2e-llm-inference-service] level: s0:c26,c20 [e2e-llm-inference-service] fsGroup: 1000690000 [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: scheduler-ha-replicas-test-epp-sa-dockercfg-mp5wd [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] tolerations: [e2e-llm-inference-service] - key: node.kubernetes.io/not-ready [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/unreachable [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoExecute [e2e-llm-inference-service] tolerationSeconds: 300 [e2e-llm-inference-service] - key: node.kubernetes.io/memory-pressure [e2e-llm-inference-service] operator: Exists [e2e-llm-inference-service] effect: NoSchedule [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] enableServiceLinks: true [e2e-llm-inference-service] preemptionPolicy: PreemptLowerPriority [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] phase: Running [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: PodReadyToStartContainers [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] - type: Initialized [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] - type: Ready [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: ContainersReady [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] reason: ContainersNotReady [e2e-llm-inference-service] message: 'containers with unready status: [main]' [e2e-llm-inference-service] - type: PodScheduled [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastProbeTime: null [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] hostIP: 10.0.138.159 [e2e-llm-inference-service] hostIPs: [e2e-llm-inference-service] - ip: 10.0.138.159 [e2e-llm-inference-service] podIP: 10.134.0.57 [e2e-llm-inference-service] podIPs: [e2e-llm-inference-service] - ip: 10.134.0.57 [e2e-llm-inference-service] startTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] containerStatuses: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] state: [e2e-llm-inference-service] running: [e2e-llm-inference-service] startedAt: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] lastState: {} [e2e-llm-inference-service] ready: false [e2e-llm-inference-service] restartCount: 0 [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] imageID: ghcr.io/llm-d/llm-d-router-endpoint-picker@sha256:06b6c75d77afd0e07053402752a9736c2dfbc12a306d0d37d963aac4c1d4e6a6 [e2e-llm-inference-service] containerID: cri-o://9063cb3f628d0cf165311b58f86176528778655773c38daaf561d6c7df0e38b0 [e2e-llm-inference-service] started: true [e2e-llm-inference-service] allocatedResources: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] - name: kube-api-access-k87xh [e2e-llm-inference-service] mountPath: /var/run/secrets/kubernetes.io/serviceaccount [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] recursiveReadOnly: Disabled [e2e-llm-inference-service] user: [e2e-llm-inference-service] linux: [e2e-llm-inference-service] uid: 1000690000 [e2e-llm-inference-service] gid: 0 [e2e-llm-inference-service] supplementalGroups: [e2e-llm-inference-service] - 0 [e2e-llm-inference-service] - 1000690000 [e2e-llm-inference-service] qosClass: Burstable [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 6f57c442-7838-439a-9630-5b934e07fbfd [e2e-llm-inference-service] resourceVersion: '92069' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] openshift.io/internal-registry-pull-secret-ref: scheduler-ha-replicas-test-epp-sa-dockercfg-mp5wd [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: openshift.io/image-registry-pull-secrets_service-account-controller [e2e-llm-inference-service] operation: Apply [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:imagePullSecrets: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] f:openshift.io/internal-registry-pull-secret-ref: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] k:{"name":"scheduler-ha-replicas-test-epp-sa-dockercfg-mp5wd"}: {} [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:secrets: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"default-dockercfg-x9l7t"}: {} [e2e-llm-inference-service] k:{"name":"seaweedfs-s3-creds"}: {} [e2e-llm-inference-service] secrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: seaweedfs-s3-creds [e2e-llm-inference-service] - name: scheduler-ha-replicas-test-epp-sa-dockercfg-mp5wd [e2e-llm-inference-service] imagePullSecrets: [e2e-llm-inference-service] - name: default-dockercfg-x9l7t [e2e-llm-inference-service] - name: scheduler-ha-replicas-test-epp-sa-dockercfg-mp5wd [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: ServiceAccount [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-service [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 7986b1f6-6c14-48a7-9c71-197c6c5a71bb [e2e-llm-inference-service] resourceVersion: '92087' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] k:{"port":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] targetPort: grpc [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] targetPort: grpc-health [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] targetPort: metrics [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] targetPort: zmq [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] clusterIP: 172.31.90.121 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.90.121 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 4507bf93-4567-40a9-824a-4e05fa33982b [e2e-llm-inference-service] resourceVersion: '92055' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:internalTrafficPolicy: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"port":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:targetPort: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:sessionAffinity: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] targetPort: 8000 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] clusterIP: 172.31.80.52 [e2e-llm-inference-service] clusterIPs: [e2e-llm-inference-service] - 172.31.80.52 [e2e-llm-inference-service] type: ClusterIP [e2e-llm-inference-service] sessionAffinity: None [e2e-llm-inference-service] ipFamilies: [e2e-llm-inference-service] - IPv4 [e2e-llm-inference-service] ipFamilyPolicy: SingleStack [e2e-llm-inference-service] internalTrafficPolicy: Cluster [e2e-llm-inference-service] status: [e2e-llm-inference-service] loadBalancer: {} [e2e-llm-inference-service] apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 6d252dd3-f2a8-4bed-aa86-6020aead6288 [e2e-llm-inference-service] resourceVersion: '92066' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:unavailableReplicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:rollingUpdate: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:maxSurge: {} [e2e-llm-inference-service] f:maxUnavailable: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-ha-replicas-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/llm-d-inference-sim [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --model [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --mode [e2e-llm-inference-service] - random [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: RollingUpdate [e2e-llm-inference-service] rollingUpdate: [e2e-llm-inference-service] maxUnavailable: 25% [e2e-llm-inference-service] maxSurge: 25% [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] updatedReplicas: 1 [e2e-llm-inference-service] unavailableReplicas: 1 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] reason: MinimumReplicasUnavailable [e2e-llm-inference-service] message: Deployment does not have minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] reason: ReplicaSetUpdated [e2e-llm-inference-service] message: ReplicaSet "scheduler-ha-replicas-test-kserve-5c5cfcf776" is progressing. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: c0720d1c-5991-4dcc-ae98-d704f5aeef1f [e2e-llm-inference-service] resourceVersion: '92108' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Available"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Progressing"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:lastUpdateTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:unavailableReplicas: {} [e2e-llm-inference-service] f:updatedReplicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:progressDeadlineSeconds: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:revisionHistoryLimit: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:strategy: [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 2 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-ha-replicas-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --ha-enable-leader-election [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n-\ [e2e-llm-inference-service] \ name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] serviceAccount: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] strategy: [e2e-llm-inference-service] type: Recreate [e2e-llm-inference-service] revisionHistoryLimit: 10 [e2e-llm-inference-service] progressDeadlineSeconds: 600 [e2e-llm-inference-service] status: [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] replicas: 2 [e2e-llm-inference-service] updatedReplicas: 2 [e2e-llm-inference-service] unavailableReplicas: 2 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - type: Available [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] reason: MinimumReplicasUnavailable [e2e-llm-inference-service] message: Deployment does not have minimum availability. [e2e-llm-inference-service] - type: Progressing [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] lastUpdateTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] reason: ReplicaSetUpdated [e2e-llm-inference-service] message: ReplicaSet "scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dbc4" [e2e-llm-inference-service] is progressing. [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-5c5cfcf776 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 0e32e010-9a87-43e0-89ce-35a224a99541 [e2e-llm-inference-service] resourceVersion: '92065' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 5c5cfcf776 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '1' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '2' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve [e2e-llm-inference-service] uid: 6d252dd3-f2a8-4bed-aa86-6020aead6288 [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"6d252dd3-f2a8-4bed-aa86-6020aead6288"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:llm-d.ai/role: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"HF_HUB_CACHE"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"HOME"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] k:{"name":"VLLM_LOGGING_LEVEL"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":8000,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:limits: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:startupProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:httpGet: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:path: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:scheme: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/dev/shm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/models"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"dshm"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:medium: {} [e2e-llm-inference-service] f:sizeLimit: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"home"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"model-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tmp-dir"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 5c5cfcf776 [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] llm-d.ai/role: both [e2e-llm-inference-service] pod-template-hash: 5c5cfcf776 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] emptyDir: [e2e-llm-inference-service] medium: Memory [e2e-llm-inference-service] sizeLimit: 1Gi [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-ha-replicas-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-inference-sim:v0.8.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/llm-d-inference-sim [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --port [e2e-llm-inference-service] - '8000' [e2e-llm-inference-service] - --model [e2e-llm-inference-service] - facebook/opt-125m [e2e-llm-inference-service] - --mode [e2e-llm-inference-service] - random [e2e-llm-inference-service] - --ssl-certfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.crt [e2e-llm-inference-service] - --ssl-keyfile [e2e-llm-inference-service] - /var/run/kserve/tls/tls.key [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - containerPort: 8000 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: HOME [e2e-llm-inference-service] value: /home [e2e-llm-inference-service] - name: VLLM_LOGGING_LEVEL [e2e-llm-inference-service] value: INFO [e2e-llm-inference-service] - name: HF_HUB_CACHE [e2e-llm-inference-service] value: /models [e2e-llm-inference-service] resources: [e2e-llm-inference-service] limits: [e2e-llm-inference-service] cpu: '1' [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 200m [e2e-llm-inference-service] memory: 2Gi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: home [e2e-llm-inference-service] mountPath: /home [e2e-llm-inference-service] - name: tmp-dir [e2e-llm-inference-service] mountPath: /tmp [e2e-llm-inference-service] - name: dshm [e2e-llm-inference-service] mountPath: /dev/shm [e2e-llm-inference-service] - name: model-cache [e2e-llm-inference-service] mountPath: /models [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 10 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 1 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 2 [e2e-llm-inference-service] startupProbe: [e2e-llm-inference-service] httpGet: [e2e-llm-inference-service] path: /health [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] scheme: HTTPS [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 60 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 1 [e2e-llm-inference-service] fullyLabeledReplicas: 1 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dbc4 [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: b92c1a65-aee7-4258-8de7-ecd6bc2982cc [e2e-llm-inference-service] resourceVersion: '92107' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 6f5f85dbc4 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] deployment.kubernetes.io/desired-replicas: '2' [e2e-llm-inference-service] deployment.kubernetes.io/max-replicas: '2' [e2e-llm-inference-service] deployment.kubernetes.io/revision: '1' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: apps/v1 [e2e-llm-inference-service] kind: Deployment [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler [e2e-llm-inference-service] uid: c0720d1c-5991-4dcc-ae98-d704f5aeef1f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/desired-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/max-replicas: {} [e2e-llm-inference-service] f:deployment.kubernetes.io/revision: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"c0720d1c-5991-4dcc-ae98-d704f5aeef1f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] f:selector: {} [e2e-llm-inference-service] f:template: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/version: {} [e2e-llm-inference-service] f:certificates.kserve.io/expiration-v2: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:pod-template-hash: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] f:containers: [e2e-llm-inference-service] k:{"name":"main"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:args: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:env: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"SSL_CERT_DIR"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:image: {} [e2e-llm-inference-service] f:imagePullPolicy: {} [e2e-llm-inference-service] f:lifecycle: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:preStop: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:command: {} [e2e-llm-inference-service] f:livenessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:ports: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"containerPort":5557,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9002,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9003,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] k:{"containerPort":9090,"protocol":"TCP"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:containerPort: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:protocol: {} [e2e-llm-inference-service] f:readinessProbe: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureThreshold: {} [e2e-llm-inference-service] f:grpc: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:port: {} [e2e-llm-inference-service] f:service: {} [e2e-llm-inference-service] f:initialDelaySeconds: {} [e2e-llm-inference-service] f:periodSeconds: {} [e2e-llm-inference-service] f:successThreshold: {} [e2e-llm-inference-service] f:timeoutSeconds: {} [e2e-llm-inference-service] f:resources: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:requests: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:cpu: {} [e2e-llm-inference-service] f:memory: {} [e2e-llm-inference-service] f:securityContext: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:allowPrivilegeEscalation: {} [e2e-llm-inference-service] f:capabilities: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:drop: {} [e2e-llm-inference-service] f:readOnlyRootFilesystem: {} [e2e-llm-inference-service] f:runAsNonRoot: {} [e2e-llm-inference-service] f:seccompProfile: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:terminationMessagePath: {} [e2e-llm-inference-service] f:terminationMessagePolicy: {} [e2e-llm-inference-service] f:volumeMounts: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"mountPath":"/tmp/tokenizer"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"mountPath":"/var/run/kserve/tls"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:mountPath: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:readOnly: {} [e2e-llm-inference-service] f:dnsPolicy: {} [e2e-llm-inference-service] f:restartPolicy: {} [e2e-llm-inference-service] f:schedulerName: {} [e2e-llm-inference-service] f:securityContext: {} [e2e-llm-inference-service] f:serviceAccount: {} [e2e-llm-inference-service] f:serviceAccountName: {} [e2e-llm-inference-service] f:terminationGracePeriodSeconds: {} [e2e-llm-inference-service] f:volumes: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"name":"tls-certs"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:secret: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:defaultMode: {} [e2e-llm-inference-service] f:secretName: {} [e2e-llm-inference-service] k:{"name":"tokenizer-cache"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-tmp"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] k:{"name":"tokenizer-uds"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:emptyDir: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:fullyLabeledReplicas: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] f:replicas: {} [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] spec: [e2e-llm-inference-service] replicas: 2 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 6f5f85dbc4 [e2e-llm-inference-service] template: [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] pod-template-hash: 6f5f85dbc4 [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] app.kubernetes.io/version: 0.9.0 [e2e-llm-inference-service] certificates.kserve.io/expiration-v2: 'true' [e2e-llm-inference-service] spec: [e2e-llm-inference-service] volumes: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] secret: [e2e-llm-inference-service] secretName: scheduler-ha-replicas-test-kserve-self-signed-certs [e2e-llm-inference-service] defaultMode: 420 [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-tmp [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] - name: tokenizer-cache [e2e-llm-inference-service] emptyDir: {} [e2e-llm-inference-service] containers: [e2e-llm-inference-service] - name: main [e2e-llm-inference-service] image: ghcr.io/llm-d/llm-d-router-endpoint-picker:v0.9.0-rc.2 [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /app/epp [e2e-llm-inference-service] - --pool-name [e2e-llm-inference-service] - scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] - --pool-namespace [e2e-llm-inference-service] - kserve-ci-e2e-test [e2e-llm-inference-service] - --zap-encoder [e2e-llm-inference-service] - json [e2e-llm-inference-service] - --grpc-port [e2e-llm-inference-service] - '9002' [e2e-llm-inference-service] - --grpc-health-port [e2e-llm-inference-service] - '9003' [e2e-llm-inference-service] - --enable-cert-reload=true [e2e-llm-inference-service] - --secure-serving=true [e2e-llm-inference-service] - --model-server-metrics-scheme=https [e2e-llm-inference-service] - --cert-path=/var/run/kserve/tls [e2e-llm-inference-service] args: [e2e-llm-inference-service] - --ha-enable-leader-election [e2e-llm-inference-service] - --config-text [e2e-llm-inference-service] - "apiVersion: inference.networking.x-k8s.io/v1alpha1\nkind: EndpointPickerConfig\n\ [e2e-llm-inference-service] plugins:\n- type: single-profile-handler\n- type: queue-scorer\n- type:\ [e2e-llm-inference-service] \ prefix-cache-scorer\n- type: max-score-picker\nschedulingProfiles:\n-\ [e2e-llm-inference-service] \ name: default\n plugins:\n - pluginRef: queue-scorer\n weight: 2\n\ [e2e-llm-inference-service] \ - pluginRef: prefix-cache-scorer\n weight: 3\n - pluginRef: max-score-picker\n" [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] containerPort: 9002 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] containerPort: 9003 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] containerPort: 9090 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] containerPort: 5557 [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] env: [e2e-llm-inference-service] - name: SSL_CERT_DIR [e2e-llm-inference-service] value: /var/run/kserve/tls:/var/run/secrets/kubernetes.io/serviceaccount:/etc/pki/tls/certs [e2e-llm-inference-service] resources: [e2e-llm-inference-service] requests: [e2e-llm-inference-service] cpu: 256m [e2e-llm-inference-service] memory: 500Mi [e2e-llm-inference-service] volumeMounts: [e2e-llm-inference-service] - name: tls-certs [e2e-llm-inference-service] readOnly: true [e2e-llm-inference-service] mountPath: /var/run/kserve/tls [e2e-llm-inference-service] - name: tokenizer-uds [e2e-llm-inference-service] mountPath: /tmp/tokenizer [e2e-llm-inference-service] livenessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: liveness [e2e-llm-inference-service] initialDelaySeconds: 5 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] readinessProbe: [e2e-llm-inference-service] grpc: [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] service: readiness [e2e-llm-inference-service] initialDelaySeconds: 30 [e2e-llm-inference-service] timeoutSeconds: 1 [e2e-llm-inference-service] periodSeconds: 10 [e2e-llm-inference-service] successThreshold: 1 [e2e-llm-inference-service] failureThreshold: 3 [e2e-llm-inference-service] lifecycle: [e2e-llm-inference-service] preStop: [e2e-llm-inference-service] exec: [e2e-llm-inference-service] command: [e2e-llm-inference-service] - /bin/sleep [e2e-llm-inference-service] - '15' [e2e-llm-inference-service] terminationMessagePath: /dev/termination-log [e2e-llm-inference-service] terminationMessagePolicy: FallbackToLogsOnError [e2e-llm-inference-service] imagePullPolicy: IfNotPresent [e2e-llm-inference-service] securityContext: [e2e-llm-inference-service] capabilities: [e2e-llm-inference-service] drop: [e2e-llm-inference-service] - ALL [e2e-llm-inference-service] runAsNonRoot: true [e2e-llm-inference-service] readOnlyRootFilesystem: true [e2e-llm-inference-service] allowPrivilegeEscalation: false [e2e-llm-inference-service] seccompProfile: [e2e-llm-inference-service] type: RuntimeDefault [e2e-llm-inference-service] restartPolicy: Always [e2e-llm-inference-service] terminationGracePeriodSeconds: 60 [e2e-llm-inference-service] dnsPolicy: ClusterFirst [e2e-llm-inference-service] serviceAccountName: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] serviceAccount: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] securityContext: {} [e2e-llm-inference-service] schedulerName: default-scheduler [e2e-llm-inference-service] status: [e2e-llm-inference-service] replicas: 2 [e2e-llm-inference-service] fullyLabeledReplicas: 2 [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] apiVersion: apps/v1 [e2e-llm-inference-service] kind: ReplicaSet [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-rb [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: abf672de-4f0a-416c-a635-13db153e40e3 [e2e-llm-inference-service] resourceVersion: '92078' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] apiGroup: rbac.authorization.k8s.io [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-role [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-role [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 956d14ea-3004-4aa0-b71a-c1e07cedb63b [e2e-llm-inference-service] resourceVersion: '92076' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] - create [e2e-llm-inference-service] - update [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - delete [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-service-zns4f [e2e-llm-inference-service] generateName: scheduler-ha-replicas-test-epp-service- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 8a771a0f-061a-4582-9882-c4dccb6d4ea7 [e2e-llm-inference-service] resourceVersion: '92154' [e2e-llm-inference-service] generation: 3 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: scheduler-ha-replicas-test-epp-service [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-service [e2e-llm-inference-service] uid: 7986b1f6-6c14-48a7-9c71-197c6c5a71bb [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"7986b1f6-6c14-48a7-9c71-197c6c5a71bb"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.134.0.57 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: false [e2e-llm-inference-service] serving: false [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85dhvggl [e2e-llm-inference-service] uid: 204bdfda-9eed-4bad-880b-7f21471eae74 [e2e-llm-inference-service] nodeName: ip-10-0-138-159.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.132.0.69 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: false [e2e-llm-inference-service] serving: false [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-router-scheduler-6f5f85d7nl6t [e2e-llm-inference-service] uid: 463c658e-bddd-4fd5-a086-b17fc708ddce [e2e-llm-inference-service] nodeName: ip-10-0-134-1.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: grpc [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9002 [e2e-llm-inference-service] - name: grpc-health [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9003 [e2e-llm-inference-service] - name: metrics [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 9090 [e2e-llm-inference-service] - name: zmq [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 5557 [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc-8br7l [e2e-llm-inference-service] generateName: scheduler-ha-replicas-test-kserve-workload-svc- [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: e0cf4cf9-0f5d-4777-a612-0f6c160a43e6 [e2e-llm-inference-service] resourceVersion: '92146' [e2e-llm-inference-service] generation: 2 [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] endpointslice.kubernetes.io/managed-by: endpointslice-controller.k8s.io [e2e-llm-inference-service] kubernetes.io/service-name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] annotations: [e2e-llm-inference-service] endpoints.kubernetes.io/last-change-trigger-time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: v1 [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] uid: 4507bf93-4567-40a9-824a-4e05fa33982b [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: kube-controller-manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:addressType: {} [e2e-llm-inference-service] f:endpoints: {} [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:annotations: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:endpoints.kubernetes.io/last-change-trigger-time: {} [e2e-llm-inference-service] f:generateName: {} [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:endpointslice.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:kubernetes.io/service-name: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"4507bf93-4567-40a9-824a-4e05fa33982b"}: {} [e2e-llm-inference-service] f:ports: {} [e2e-llm-inference-service] addressType: IPv4 [e2e-llm-inference-service] endpoints: [e2e-llm-inference-service] - addresses: [e2e-llm-inference-service] - 10.133.0.52 [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] ready: false [e2e-llm-inference-service] serving: false [e2e-llm-inference-service] terminating: false [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] kind: Pod [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-5c5cfcf776-sp8jl [e2e-llm-inference-service] uid: 0e8fac6c-4313-4304-b55c-30ffeecc7fff [e2e-llm-inference-service] nodeName: ip-10-0-141-170.ec2.internal [e2e-llm-inference-service] zone: us-east-1a [e2e-llm-inference-service] ports: [e2e-llm-inference-service] - name: https [e2e-llm-inference-service] protocol: TCP [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] appProtocol: https [e2e-llm-inference-service] apiVersion: discovery.k8s.io/v1 [e2e-llm-inference-service] kind: EndpointSlice [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-rb [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: abf672de-4f0a-416c-a635-13db153e40e3 [e2e-llm-inference-service] resourceVersion: '92078' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:roleRef: {} [e2e-llm-inference-service] f:subjects: {} [e2e-llm-inference-service] userNames: [e2e-llm-inference-service] - system:serviceaccount:kserve-ci-e2e-test:scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] groupNames: null [e2e-llm-inference-service] subjects: [e2e-llm-inference-service] - kind: ServiceAccount [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-sa [e2e-llm-inference-service] roleRef: [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-role [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: RoleBinding [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 metadata: [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-role [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] uid: 956d14ea-3004-4aa0-b71a-c1e07cedb63b [e2e-llm-inference-service] resourceVersion: '92076' [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] apiVersion: rbac.authorization.k8s.io/v1 [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:rules: {} [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - '' [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - pods [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.k8s.io [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodels [e2e-llm-inference-service] - inferenceobjectives [e2e-llm-inference-service] - inferencepools [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - inference.networking.x-k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - inferencemodelrewrites [e2e-llm-inference-service] - inferencepoolimports [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - discovery.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - endpointslices [e2e-llm-inference-service] - verbs: [e2e-llm-inference-service] - create [e2e-llm-inference-service] - delete [e2e-llm-inference-service] - get [e2e-llm-inference-service] - list [e2e-llm-inference-service] - patch [e2e-llm-inference-service] - update [e2e-llm-inference-service] - watch [e2e-llm-inference-service] attributeRestrictions: null [e2e-llm-inference-service] apiGroups: [e2e-llm-inference-service] - coordination.k8s.io [e2e-llm-inference-service] resources: [e2e-llm-inference-service] - leases [e2e-llm-inference-service] apiVersion: authorization.openshift.io/v1 [e2e-llm-inference-service] kind: Role [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:29:41Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-route [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92232' [e2e-llm-inference-service] uid: ad539bf0-cc01-46d9-8695-0e449d50fdcb [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-ha-replicas-test/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-ha-replicas-test/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-ha-replicas-test/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-ha-replicas-test/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-ha-replicas-test [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] message: 'referencing unsupported backendRef: group "inference.networking.x-k8s.io" [e2e-llm-inference-service] kind "InferencePool"' [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: InvalidKind [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:29:40Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1beta1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] - apiVersion: gateway.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] f:parents: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:29:41Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-route [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92232' [e2e-llm-inference-service] uid: ad539bf0-cc01-46d9-8695-0e449d50fdcb [e2e-llm-inference-service] spec: [e2e-llm-inference-service] parentRefs: [e2e-llm-inference-service] - group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] rules: [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-ha-replicas-test/v1/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/chat/completions [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-ha-replicas-test/v1/chat/completions [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/chat/completions/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/responses [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-ha-replicas-test/v1/responses [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/responses/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: /v1/messages [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-ha-replicas-test/v1/messages [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: inference.networking.x-k8s.io [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: /v1/messages/ [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] filters: [e2e-llm-inference-service] - type: URLRewrite [e2e-llm-inference-service] urlRewrite: [e2e-llm-inference-service] path: [e2e-llm-inference-service] replacePrefixMatch: / [e2e-llm-inference-service] type: ReplacePrefixMatch [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: /kserve-ci-e2e-test/scheduler-ha-replicas-test [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] - backendRefs: [e2e-llm-inference-service] - group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] port: 8000 [e2e-llm-inference-service] weight: 1 [e2e-llm-inference-service] matches: [e2e-llm-inference-service] - headers: [e2e-llm-inference-service] - name: X-Gateway-Model-Name [e2e-llm-inference-service] type: Exact [e2e-llm-inference-service] value: publishers/kserve-ci-e2e-test/models/facebook/opt-125m [e2e-llm-inference-service] path: [e2e-llm-inference-service] type: PathPrefix [e2e-llm-inference-service] value: / [e2e-llm-inference-service] timeouts: [e2e-llm-inference-service] backendRequest: 0s [e2e-llm-inference-service] request: 0s [e2e-llm-inference-service] status: [e2e-llm-inference-service] parents: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] message: Route was valid [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] message: 'referencing unsupported backendRef: group "inference.networking.x-k8s.io" [e2e-llm-inference-service] kind "InferencePool"' [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: InvalidKind [e2e-llm-inference-service] status: 'False' [e2e-llm-inference-service] type: ResolvedRefs [e2e-llm-inference-service] controllerName: openshift.io/gateway-controller/v1 [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:29:40Z' [e2e-llm-inference-service] message: Object affected by AuthPolicy [kserve-ci-e2e-test/scheduler-ha-replicas-test-kserve-route-authn [e2e-llm-inference-service] openshift-ingress/openshift-ai-inference-authn] [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: kuadrant.io/AuthPolicyAffected [e2e-llm-inference-service] controllerName: kuadrant.io/policy-controller [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Gateway [e2e-llm-inference-service] name: openshift-ai-inference [e2e-llm-inference-service] namespace: openshift-ingress [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:appProtocol: {} [e2e-llm-inference-service] f:endpointPickerRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:port: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:number: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:matchLabels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPorts: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] - apiVersion: inference.networking.k8s.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] manager: pilot-discovery [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92097' [e2e-llm-inference-service] uid: 6cba106f-3a4d-4d91-a591-bb727e657132 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] appProtocol: http [e2e-llm-inference-service] endpointPickerRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-service [e2e-llm-inference-service] port: [e2e-llm-inference-service] number: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] matchLabels: [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPorts: [e2e-llm-inference-service] - number: 8000 [e2e-llm-inference-service] status: {} [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] kind: AuthPolicy [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:40Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-policies [e2e-llm-inference-service] app.kubernetes.io/managed-by: odh-model-controller [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/managed-by: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:rules: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:authentication: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:public: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:anonymous: {} [e2e-llm-inference-service] f:credentials: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:overrides: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:fairness: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:value: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:response: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:success: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:headers: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:x-gateway-inference-fairness-id: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:x-gateway-inference-objective: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:metrics: {} [e2e-llm-inference-service] f:plain: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:expression: {} [e2e-llm-inference-service] f:priority: {} [e2e-llm-inference-service] f:targetRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:40Z' [e2e-llm-inference-service] - apiVersion: kuadrant.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:status: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:conditions: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"type":"Accepted"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] k:{"type":"Enforced"}: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:lastTransitionTime: {} [e2e-llm-inference-service] f:message: {} [e2e-llm-inference-service] f:reason: {} [e2e-llm-inference-service] f:status: {} [e2e-llm-inference-service] f:type: {} [e2e-llm-inference-service] f:observedGeneration: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] subresource: status [e2e-llm-inference-service] time: '2026-07-02T14:29:43Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-route-authn [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92252' [e2e-llm-inference-service] uid: ac3190d3-7c0b-4c5d-b52e-f1d143547ff1 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] rules: [e2e-llm-inference-service] authentication: [e2e-llm-inference-service] public: [e2e-llm-inference-service] anonymous: {} [e2e-llm-inference-service] credentials: {} [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] overrides: [e2e-llm-inference-service] fairness: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] objective: [e2e-llm-inference-service] value: unauthenticated [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] response: [e2e-llm-inference-service] success: [e2e-llm-inference-service] headers: [e2e-llm-inference-service] x-gateway-inference-fairness-id: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.fairness [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] x-gateway-inference-objective: [e2e-llm-inference-service] metrics: false [e2e-llm-inference-service] plain: [e2e-llm-inference-service] expression: auth.identity.objective [e2e-llm-inference-service] priority: 0 [e2e-llm-inference-service] targetRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: HTTPRoute [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-route [e2e-llm-inference-service] status: [e2e-llm-inference-service] conditions: [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:29:41Z' [e2e-llm-inference-service] message: AuthPolicy has been accepted [e2e-llm-inference-service] reason: Accepted [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] - lastTransitionTime: '2026-07-02T14:29:43Z' [e2e-llm-inference-service] message: AuthPolicy has been successfully enforced [e2e-llm-inference-service] reason: Enforced [e2e-llm-inference-service] status: 'True' [e2e-llm-inference-service] type: Enforced [e2e-llm-inference-service] observedGeneration: 1 [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92131' [e2e-llm-inference-service] uid: 7c57afe5-6474-44ec-b719-1f1c804c4f1f [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-ha-replicas-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-ha-replicas-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92149' [e2e-llm-inference-service] uid: 0ab8fa82-3ebd-44a3-ae7d-9208c0fbc23a [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-ha-replicas-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-ha-replicas-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92131' [e2e-llm-inference-service] uid: 7c57afe5-6474-44ec-b719-1f1c804c4f1f [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-ha-replicas-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-ha-replicas-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1beta1 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92149' [e2e-llm-inference-service] uid: 0ab8fa82-3ebd-44a3-ae7d-9208c0fbc23a [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-ha-replicas-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-ha-replicas-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-scheduler [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92131' [e2e-llm-inference-service] uid: 7c57afe5-6474-44ec-b719-1f1c804c4f1f [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-ha-replicas-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] insecureSkipVerify: true [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-ha-replicas-test-epp-service.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: networking.istio.io/v1alpha3 [e2e-llm-inference-service] kind: DestinationRule [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-workload [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] llm-d.ai/managed: 'true' [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: networking.istio.io/v1 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:llm-d.ai/managed: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:exportTo: {} [e2e-llm-inference-service] f:host: {} [e2e-llm-inference-service] f:trafficPolicy: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:tls: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:caCertificates: {} [e2e-llm-inference-service] f:insecureSkipVerify: {} [e2e-llm-inference-service] f:mode: {} [e2e-llm-inference-service] f:sni: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:39Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-kserve-workload-svc [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92149' [e2e-llm-inference-service] uid: 0ab8fa82-3ebd-44a3-ae7d-9208c0fbc23a [e2e-llm-inference-service] spec: [e2e-llm-inference-service] exportTo: [e2e-llm-inference-service] - '*' [e2e-llm-inference-service] host: scheduler-ha-replicas-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] trafficPolicy: [e2e-llm-inference-service] tls: [e2e-llm-inference-service] caCertificates: /var/run/secrets/kubernetes.io/serviceaccount/service-ca.crt [e2e-llm-inference-service] insecureSkipVerify: false [e2e-llm-inference-service] mode: SIMPLE [e2e-llm-inference-service] sni: scheduler-ha-replicas-test-kserve-workload-svc.kserve-ci-e2e-test.svc.cluster.local [e2e-llm-inference-service] [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1254 --- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1255 apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] kind: InferencePool [e2e-llm-inference-service] metadata: [e2e-llm-inference-service] creationTimestamp: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] generation: 1 [e2e-llm-inference-service] labels: [e2e-llm-inference-service] app.kubernetes.io/component: llminferenceservice-router-scheduler [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] managedFields: [e2e-llm-inference-service] - apiVersion: inference.networking.x-k8s.io/v1alpha2 [e2e-llm-inference-service] fieldsType: FieldsV1 [e2e-llm-inference-service] fieldsV1: [e2e-llm-inference-service] f:metadata: [e2e-llm-inference-service] f:labels: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/component: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:ownerReferences: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] k:{"uid":"132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f"}: {} [e2e-llm-inference-service] f:spec: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:extensionRef: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:failureMode: {} [e2e-llm-inference-service] f:group: {} [e2e-llm-inference-service] f:kind: {} [e2e-llm-inference-service] f:name: {} [e2e-llm-inference-service] f:portNumber: {} [e2e-llm-inference-service] f:selector: [e2e-llm-inference-service] .: {} [e2e-llm-inference-service] f:app.kubernetes.io/name: {} [e2e-llm-inference-service] f:app.kubernetes.io/part-of: {} [e2e-llm-inference-service] f:kserve.io/component: {} [e2e-llm-inference-service] f:targetPortNumber: {} [e2e-llm-inference-service] manager: manager [e2e-llm-inference-service] operation: Update [e2e-llm-inference-service] time: '2026-07-02T14:29:38Z' [e2e-llm-inference-service] name: scheduler-ha-replicas-test-inference-pool [e2e-llm-inference-service] namespace: kserve-ci-e2e-test [e2e-llm-inference-service] ownerReferences: [e2e-llm-inference-service] - apiVersion: serving.kserve.io/v1alpha2 [e2e-llm-inference-service] blockOwnerDeletion: true [e2e-llm-inference-service] controller: true [e2e-llm-inference-service] kind: LLMInferenceService [e2e-llm-inference-service] name: scheduler-ha-replicas-test [e2e-llm-inference-service] uid: 132d7ebf-1b2e-4603-9f6d-c2cd0b01d53f [e2e-llm-inference-service] resourceVersion: '92098' [e2e-llm-inference-service] uid: 4fa3ce31-d712-407e-886c-c5f66961aa47 [e2e-llm-inference-service] spec: [e2e-llm-inference-service] extensionRef: [e2e-llm-inference-service] failureMode: FailOpen [e2e-llm-inference-service] group: '' [e2e-llm-inference-service] kind: Service [e2e-llm-inference-service] name: scheduler-ha-replicas-test-epp-service [e2e-llm-inference-service] portNumber: 9002 [e2e-llm-inference-service] selector: [e2e-llm-inference-service] app.kubernetes.io/name: scheduler-ha-replicas-test [e2e-llm-inference-service] app.kubernetes.io/part-of: llminferenceservice [e2e-llm-inference-service] kserve.io/component: workload [e2e-llm-inference-service] targetPortNumber: 8000 [e2e-llm-inference-service] status: [e2e-llm-inference-service] parent: [e2e-llm-inference-service] - conditions: [e2e-llm-inference-service] - lastTransitionTime: '1970-01-01T00:00:00Z' [e2e-llm-inference-service] message: Waiting for controller [e2e-llm-inference-service] reason: Pending [e2e-llm-inference-service] status: Unknown [e2e-llm-inference-service] type: Accepted [e2e-llm-inference-service] parentRef: [e2e-llm-inference-service] group: gateway.networking.k8s.io [e2e-llm-inference-service] kind: Status [e2e-llm-inference-service] name: default [e2e-llm-inference-service] [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_llm_inference_service] [2026-07-02T14:29:47.049585] end - ❌ 11.828s: ❌ Exception when calling CustomObjectsApi->get_namespaced_custom_object for LLMInferenceService: (500) [e2e-llm-inference-service] Reason: Internal Server Error [e2e-llm-inference-service] HTTP response headers: HTTPHeaderDict({'Audit-Id': 'e505316d-488e-4927-80b5-377bc18852df', 'Cache-Control': 'no-cache, private', 'Content-Type': 'application/json', 'Strict-Transport-Security': 'max-age=31536000; includeSubDomains; preload', 'X-Kubernetes-Pf-Flowschema-Uid': 'a8664215-74a2-48d8-b231-e657e37b3200', 'X-Kubernetes-Pf-Prioritylevel-Uid': '0e7454c5-cfcf-4e4f-b47e-e88e460e3769', 'Date': 'Thu, 02 Jul 2026 14:29:45 GMT', 'Content-Length': '264'}) [e2e-llm-inference-service] HTTP response body: {"kind":"Status","apiVersion":"v1","metadata":{},"status":"Failure","message":"conversion webhook for serving.kserve.io/v1alpha2, Kind=LLMInferenceService failed: Post \"https://llmisvc-webhook-server-service.kserve.svc:443/convert?timeout=30s\": EOF","code":500} [e2e-llm-inference-service] ___ test_prestop_hook[router-managed-workload-single-cpu-model-fb-opt-125m] ____ [e2e-llm-inference-service] [gw1] linux -- Python 3.11.13 /workspace/source/python/kserve/.venv/bin/python [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def _new_conn(self) -> socket.socket: [e2e-llm-inference-service] """Establish a socket connection and set nodelay settings on it. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: New socket connection. [e2e-llm-inference-service] """ [e2e-llm-inference-service] try: [e2e-llm-inference-service] > sock = connection.create_connection( [e2e-llm-inference-service] (self._dns_host, self.port), [e2e-llm-inference-service] self.timeout, [e2e-llm-inference-service] source_address=self.source_address, [e2e-llm-inference-service] socket_options=self.socket_options, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:204: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] address = ('a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', 6443) [e2e-llm-inference-service] timeout = None, source_address = None, socket_options = [(6, 1, 1)] [e2e-llm-inference-service] [e2e-llm-inference-service] def create_connection( [e2e-llm-inference-service] address: tuple[str, int], [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] source_address: tuple[str, int] | None = None, [e2e-llm-inference-service] socket_options: _TYPE_SOCKET_OPTIONS | None = None, [e2e-llm-inference-service] ) -> socket.socket: [e2e-llm-inference-service] """Connect to *address* and return the socket object. [e2e-llm-inference-service] [e2e-llm-inference-service] Convenience function. Connect to *address* (a 2-tuple ``(host, [e2e-llm-inference-service] port)``) and return the socket object. Passing the optional [e2e-llm-inference-service] *timeout* parameter will set the timeout on the socket instance [e2e-llm-inference-service] before attempting to connect. If no *timeout* is supplied, the [e2e-llm-inference-service] global default timeout setting returned by :func:`socket.getdefaulttimeout` [e2e-llm-inference-service] is used. If *source_address* is set it must be a tuple of (host, port) [e2e-llm-inference-service] for the socket to bind as a source address before making the connection. [e2e-llm-inference-service] An host of '' or port 0 tells the OS to use the default. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] host, port = address [e2e-llm-inference-service] if host.startswith("["): [e2e-llm-inference-service] host = host.strip("[]") [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Using the value from allowed_gai_family() in the context of getaddrinfo lets [e2e-llm-inference-service] # us select whether to work with IPv4 DNS records, IPv6 records, or both. [e2e-llm-inference-service] # The original create_connection function always returns all records. [e2e-llm-inference-service] family = allowed_gai_family() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] host.encode("idna") [e2e-llm-inference-service] except UnicodeError: [e2e-llm-inference-service] raise LocationParseError(f"'{host}', label empty or too long") from None [e2e-llm-inference-service] [e2e-llm-inference-service] > for res in socket.getaddrinfo(host, port, family, socket.SOCK_STREAM): [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/util/connection.py:60: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] host = 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' [e2e-llm-inference-service] port = 6443, family = [e2e-llm-inference-service] type = , proto = 0, flags = 0 [e2e-llm-inference-service] [e2e-llm-inference-service] def getaddrinfo(host, port, family=0, type=0, proto=0, flags=0): [e2e-llm-inference-service] """Resolve host and port into list of address info entries. [e2e-llm-inference-service] [e2e-llm-inference-service] Translate the host/port argument into a sequence of 5-tuples that contain [e2e-llm-inference-service] all the necessary arguments for creating a socket connected to that service. [e2e-llm-inference-service] host is a domain name, a string representation of an IPv4/v6 address or [e2e-llm-inference-service] None. port is a string service name such as 'http', a numeric port number or [e2e-llm-inference-service] None. By passing None as the value of host and port, you can pass NULL to [e2e-llm-inference-service] the underlying C API. [e2e-llm-inference-service] [e2e-llm-inference-service] The family, type and proto arguments can be optionally specified in order to [e2e-llm-inference-service] narrow the list of addresses returned. Passing zero as a value for each of [e2e-llm-inference-service] these arguments selects the full range of results. [e2e-llm-inference-service] """ [e2e-llm-inference-service] # We override this function since we want to translate the numeric family [e2e-llm-inference-service] # and socket type values to enum constants. [e2e-llm-inference-service] addrlist = [] [e2e-llm-inference-service] > for res in _socket.getaddrinfo(host, port, family, type, proto, flags): [e2e-llm-inference-service] E socket.gaierror: [Errno -2] Name or service not known [e2e-llm-inference-service] [e2e-llm-inference-service] /usr/lib64/python3.11/socket.py:974: gaierror [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False, timeout = None, pool_timeout = None [e2e-llm-inference-service] release_conn = True, chunked = False, body_pos = None, preload_content = True [e2e-llm-inference-service] decode_content = True, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/api/v1/namespaces/kserve-ci-e2e-test/pods', query='labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload', fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] > response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:787: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=None, read=None, total=None), chunked = False [e2e-llm-inference-service] response_conn = None, preload_content = True, decode_content = True [e2e-llm-inference-service] enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._validate_conn(conn) [e2e-llm-inference-service] except (SocketTimeout, BaseSSLError) as e: [e2e-llm-inference-service] self._raise_timeout(err=e, url=url, timeout_value=conn.timeout) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # _validate_conn() starts the connection to an HTTPS proxy [e2e-llm-inference-service] # so we need to wrap errors with 'ProxyError' here too. [e2e-llm-inference-service] except ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] # If the connection didn't successfully connect to it's proxy [e2e-llm-inference-service] # then there [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, (OSError, NewConnectionError, TimeoutError, SSLError) [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] > raise new_e [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:488: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] timeout = Timeout(connect=None, read=None, total=None), chunked = False [e2e-llm-inference-service] response_conn = None, preload_content = True, decode_content = True [e2e-llm-inference-service] enforce_content_length = True [e2e-llm-inference-service] [e2e-llm-inference-service] def _make_request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] conn: BaseHTTPConnection, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | None = None, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] response_conn: BaseHTTPConnection | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] enforce_content_length: bool = True, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Perform a request on a given urllib connection object taken from our [e2e-llm-inference-service] pool. [e2e-llm-inference-service] [e2e-llm-inference-service] :param conn: [e2e-llm-inference-service] a connection from one of our connection pools [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] Pass ``None`` to retry until you receive a response. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response_conn: [e2e-llm-inference-service] Set this to ``None`` if you will handle releasing the connection or [e2e-llm-inference-service] set the connection to have the response release it. [e2e-llm-inference-service] [e2e-llm-inference-service] :param preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded during construction. [e2e-llm-inference-service] [e2e-llm-inference-service] :param decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param enforce_content_length: [e2e-llm-inference-service] Enforce content length checking. Body returned by server must match [e2e-llm-inference-service] value of Content-Length header, if present. Otherwise, raise error. [e2e-llm-inference-service] """ [e2e-llm-inference-service] self.num_requests += 1 [e2e-llm-inference-service] [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] timeout_obj.start_connect() [e2e-llm-inference-service] conn.timeout = Timeout.resolve_default_timeout(timeout_obj.connect_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Trigger any extra validation we need to do. [e2e-llm-inference-service] try: [e2e-llm-inference-service] > self._validate_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:464: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] conn = [e2e-llm-inference-service] [e2e-llm-inference-service] def _validate_conn(self, conn: BaseHTTPConnection) -> None: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Called right before a request is made, after the socket is created. [e2e-llm-inference-service] """ [e2e-llm-inference-service] super()._validate_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] # Force connect early to allow us to validate the connection. [e2e-llm-inference-service] if conn.is_closed: [e2e-llm-inference-service] > conn.connect() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:1093: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def connect(self) -> None: [e2e-llm-inference-service] # Today we don't need to be doing this step before the /actual/ socket [e2e-llm-inference-service] # connection, however in the future we'll need to decide whether to [e2e-llm-inference-service] # create a new socket or re-use an existing "shared" socket as a part [e2e-llm-inference-service] # of the HTTP/2 handshake dance. [e2e-llm-inference-service] if self._tunnel_host is not None and self._tunnel_port is not None: [e2e-llm-inference-service] probe_http2_host = self._tunnel_host [e2e-llm-inference-service] probe_http2_port = self._tunnel_port [e2e-llm-inference-service] else: [e2e-llm-inference-service] probe_http2_host = self.host [e2e-llm-inference-service] probe_http2_port = self.port [e2e-llm-inference-service] [e2e-llm-inference-service] # Check if the target origin supports HTTP/2. [e2e-llm-inference-service] # If the value comes back as 'None' it means that the current thread [e2e-llm-inference-service] # is probing for HTTP/2 support. Otherwise, we're waiting for another [e2e-llm-inference-service] # probe to complete, or we get a value right away. [e2e-llm-inference-service] target_supports_http2: bool | None [e2e-llm-inference-service] if "h2" in ssl_.ALPN_PROTOCOLS: [e2e-llm-inference-service] target_supports_http2 = http2_probe.acquire_and_get( [e2e-llm-inference-service] host=probe_http2_host, port=probe_http2_port [e2e-llm-inference-service] ) [e2e-llm-inference-service] else: [e2e-llm-inference-service] # If HTTP/2 isn't going to be offered it doesn't matter if [e2e-llm-inference-service] # the target supports HTTP/2. Don't want to make a probe. [e2e-llm-inference-service] target_supports_http2 = False [e2e-llm-inference-service] [e2e-llm-inference-service] if self._connect_callback is not None: [e2e-llm-inference-service] self._connect_callback( [e2e-llm-inference-service] "before connect", [e2e-llm-inference-service] thread_id=threading.get_ident(), [e2e-llm-inference-service] target_supports_http2=target_supports_http2, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] sock: socket.socket | ssl.SSLSocket [e2e-llm-inference-service] > self.sock = sock = self._new_conn() [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:759: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] [e2e-llm-inference-service] def _new_conn(self) -> socket.socket: [e2e-llm-inference-service] """Establish a socket connection and set nodelay settings on it. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: New socket connection. [e2e-llm-inference-service] """ [e2e-llm-inference-service] try: [e2e-llm-inference-service] sock = connection.create_connection( [e2e-llm-inference-service] (self._dns_host, self.port), [e2e-llm-inference-service] self.timeout, [e2e-llm-inference-service] source_address=self.source_address, [e2e-llm-inference-service] socket_options=self.socket_options, [e2e-llm-inference-service] ) [e2e-llm-inference-service] except socket.gaierror as e: [e2e-llm-inference-service] > raise NameResolutionError(self.host, self, e) from e [e2e-llm-inference-service] E urllib3.exceptions.NameResolutionError: HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Failed to resolve 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connection.py:211: NameResolutionError [e2e-llm-inference-service] [e2e-llm-inference-service] The above exception was the direct cause of the following exception: [e2e-llm-inference-service] [e2e-llm-inference-service] test_case = TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name=... {'name': 'model-fb-opt-125m-prestop-hook-584a80fb'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m') [e2e-llm-inference-service] [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] @pytest.mark.parametrize( [e2e-llm-inference-service] "test_case", [e2e-llm-inference-service] [ [e2e-llm-inference-service] pytest.param( [e2e-llm-inference-service] TestCase( [e2e-llm-inference-service] base_refs=[ [e2e-llm-inference-service] "router-managed", [e2e-llm-inference-service] "workload-single-cpu", [e2e-llm-inference-service] "model-fb-opt-125m", [e2e-llm-inference-service] ], [e2e-llm-inference-service] prompt="KServe is a", [e2e-llm-inference-service] service_name="prestop-hook-test", [e2e-llm-inference-service] ), [e2e-llm-inference-service] marks=[pytest.mark.cluster_cpu, pytest.mark.cluster_single_node], [e2e-llm-inference-service] ), [e2e-llm-inference-service] ], [e2e-llm-inference-service] indirect=["test_case"], [e2e-llm-inference-service] ids=generate_test_id, [e2e-llm-inference-service] ) [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def test_prestop_hook(test_case: TestCase): [e2e-llm-inference-service] """Verify the preStop lifecycle hook delays pod termination by at least 15 seconds.""" [e2e-llm-inference-service] inject_k8s_proxy() [e2e-llm-inference-service] [e2e-llm-inference-service] kserve_client = KServeClient( [e2e-llm-inference-service] config_file=os.environ.get("KUBECONFIG", "~/.kube/config"), [e2e-llm-inference-service] client_configuration=client.Configuration(), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] service_name = test_case.llm_service.metadata.name [e2e-llm-inference-service] namespace = test_case.llm_service.metadata.namespace [e2e-llm-inference-service] test_failed = False [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] print(f"Creating LLMInferenceService {service_name}") [e2e-llm-inference-service] create_llmisvc(kserve_client, test_case.llm_service) [e2e-llm-inference-service] [e2e-llm-inference-service] > workload_pod_name = wait_for_workload_pod_running( [e2e-llm-inference-service] namespace, service_name, timeout_seconds=test_case.wait_timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_prestop_hook.py:79: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] args = ('kserve-ci-e2e-test', 'prestop-hook-test') [e2e-llm-inference-service] kwargs = {'timeout_seconds': 900}, func_name = 'wait_for_workload_pod_running' [e2e-llm-inference-service] timestamp_start = '2026-07-02T14:28:55.079274', start_time = 1783002535.0793319 [e2e-llm-inference-service] duration = 789.8006458282471, timestamp_end = '2026-07-02T14:42:04.879981' [e2e-llm-inference-service] [e2e-llm-inference-service] @functools.wraps(func) [e2e-llm-inference-service] def wrapper(*args, **kwargs): [e2e-llm-inference-service] func_name = func.__name__ [e2e-llm-inference-service] [e2e-llm-inference-service] timestamp_start = datetime.now().isoformat() [e2e-llm-inference-service] logger.info( [e2e-llm-inference-service] f"[{func_name}] [{timestamp_start}] start - args={args}, kwargs={kwargs}" [e2e-llm-inference-service] ) [e2e-llm-inference-service] start_time = time.time() [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] > result = func(*args, **kwargs) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/logging.py:40: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test', service_name = 'prestop-hook-test' [e2e-llm-inference-service] timeout_seconds = 900 [e2e-llm-inference-service] [e2e-llm-inference-service] @log_execution [e2e-llm-inference-service] def wait_for_workload_pod_running( [e2e-llm-inference-service] namespace: str, [e2e-llm-inference-service] service_name: str, [e2e-llm-inference-service] timeout_seconds: int = 900, [e2e-llm-inference-service] ) -> str: [e2e-llm-inference-service] """Wait for the main workload pod to reach Running state, return its name.""" [e2e-llm-inference-service] core_v1 = client.CoreV1Api() [e2e-llm-inference-service] label_selector = ( [e2e-llm-inference-service] f"app.kubernetes.io/name={service_name},kserve.io/component=workload" [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_pod_running() -> str: [e2e-llm-inference-service] pods = core_v1.list_namespaced_pod(namespace, label_selector=label_selector) [e2e-llm-inference-service] running = [ [e2e-llm-inference-service] p [e2e-llm-inference-service] for p in pods.items [e2e-llm-inference-service] if p.status.phase == "Running" [e2e-llm-inference-service] and all(cs.ready for cs in (p.status.container_statuses or [])) [e2e-llm-inference-service] ] [e2e-llm-inference-service] assert running, f"No Running pods found for {label_selector} in {namespace}" [e2e-llm-inference-service] return running[0].metadata.name [e2e-llm-inference-service] [e2e-llm-inference-service] > return wait_for(assert_pod_running, timeout=timeout_seconds, interval=5.0) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_prestop_hook.py:149: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] assertion_fn = .assert_pod_running at 0x7f3d99e06f20> [e2e-llm-inference-service] timeout = 900, interval = 5.0 [e2e-llm-inference-service] [e2e-llm-inference-service] def wait_for( [e2e-llm-inference-service] assertion_fn: Callable[[], Any], timeout: float = 5.0, interval: float = 0.1 [e2e-llm-inference-service] ) -> Any: [e2e-llm-inference-service] """Wait for the assertion to succeed within timeout.""" [e2e-llm-inference-service] deadline = time.time() + timeout [e2e-llm-inference-service] last_msg = None [e2e-llm-inference-service] while True: [e2e-llm-inference-service] try: [e2e-llm-inference-service] > return assertion_fn() [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:1215: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] def assert_pod_running() -> str: [e2e-llm-inference-service] > pods = core_v1.list_namespaced_pod(namespace, label_selector=label_selector) [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_prestop_hook.py:139: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test' [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True, 'label_selector': 'app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload'} [e2e-llm-inference-service] [e2e-llm-inference-service] def list_namespaced_pod(self, namespace, **kwargs): # noqa: E501 [e2e-llm-inference-service] """list_namespaced_pod # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] list or watch objects of kind Pod # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.list_namespaced_pod(namespace, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str namespace: object name and auth scope, such as for teams and projects (required) [e2e-llm-inference-service] :param str pretty: If 'true', then the output is pretty printed. Defaults to 'false' unless the user-agent indicates a browser or command-line HTTP tool (curl and wget). [e2e-llm-inference-service] :param bool allow_watch_bookmarks: allowWatchBookmarks requests watch events with type \"BOOKMARK\". Servers that do not implement bookmarks may ignore this flag and bookmarks are sent at the server's discretion. Clients should not assume bookmarks are returned at any specific interval, nor may they assume the server will send any BOOKMARK event during a session. If this is not a watch, this field is ignored. [e2e-llm-inference-service] :param str _continue: The continue option should be set when retrieving more results from the server. Since this value is server defined, clients may only use the continue value from a previous query result with identical query parameters (except for the value of continue) and the server may reject a continue value it does not recognize. If the specified continue value is no longer valid whether due to expiration (generally five to fifteen minutes) or a configuration change on the server, the server will respond with a 410 ResourceExpired error together with a continue token. If the client needs a consistent list, it must restart their list without the continue field. Otherwise, the client may send another list request with the token received with the 410 error, the server will respond with a list starting from the next key, but from the latest snapshot, which is inconsistent from the previous list results - objects that are created, modified, or deleted after the first list request will be included in the response, as long as their keys are after the \"next key\". This field is not supported when watch is true. Clients may start a watch from the last resourceVersion value returned by the server and not miss any modifications. [e2e-llm-inference-service] :param str field_selector: A selector to restrict the list of returned objects by their fields. Defaults to everything. [e2e-llm-inference-service] :param str label_selector: A selector to restrict the list of returned objects by their labels. Defaults to everything. [e2e-llm-inference-service] :param int limit: limit is a maximum number of responses to return for a list call. If more items exist, the server will set the `continue` field on the list metadata to a value that can be used with the same initial query to retrieve the next set of results. Setting a limit may return fewer than the requested amount of items (up to zero items) in the event all requested objects are filtered out and clients should only use the presence of the continue field to determine whether more results are available. Servers may choose not to support the limit argument and will return all of the available results. If limit is specified and the continue field is empty, clients may assume that no more results are available. This field is not supported if watch is true. The server guarantees that the objects returned when using continue will be identical to issuing a single list call without a limit - that is, no objects created, modified, or deleted after the first request is issued will be included in any subsequent continued requests. This is sometimes referred to as a consistent snapshot, and ensures that a client that is using limit to receive smaller chunks of a very large result can ensure they see all possible objects. If objects are updated during a chunked list the version of the object that was present at the time the first list result was calculated is returned. [e2e-llm-inference-service] :param str resource_version: resourceVersion sets a constraint on what resource versions a request may be served from. See https://kubernetes.io/docs/reference/using-api/api-concepts/#resource-versions for details. Defaults to unset [e2e-llm-inference-service] :param str resource_version_match: resourceVersionMatch determines how resourceVersion is applied to list calls. It is highly recommended that resourceVersionMatch be set for list calls where resourceVersion is set See https://kubernetes.io/docs/reference/using-api/api-concepts/#resource-versions for details. Defaults to unset [e2e-llm-inference-service] :param bool send_initial_events: `sendInitialEvents=true` may be set together with `watch=true`. In that case, the watch stream will begin with synthetic events to produce the current state of objects in the collection. Once all such events have been sent, a synthetic \"Bookmark\" event will be sent. The bookmark will report the ResourceVersion (RV) corresponding to the set of objects, and be marked with `\"k8s.io/initial-events-end\": \"true\"` annotation. Afterwards, the watch stream will proceed as usual, sending watch events corresponding to changes (subsequent to the RV) to objects watched. When `sendInitialEvents` option is set, we require `resourceVersionMatch` option to also be set. The semantic of the watch request is as following: - `resourceVersionMatch` = NotOlderThan is interpreted as \"data at least as new as the provided `resourceVersion`\" and the bookmark event is send when the state is synced to a `resourceVersion` at least as fresh as the one provided by the ListOptions. If `resourceVersion` is unset, this is interpreted as \"consistent read\" and the bookmark event is send when the state is synced at least to the moment when request started being processed. - `resourceVersionMatch` set to any other value or unset Invalid error is returned. Defaults to true if `resourceVersion=\"\"` or `resourceVersion=\"0\"` (for backward compatibility reasons) and to false otherwise. [e2e-llm-inference-service] :param int timeout_seconds: Timeout for the list/watch call. This limits the duration of the call, regardless of any activity or inactivity. [e2e-llm-inference-service] :param bool watch: Watch for changes to the described resources and return them as a stream of add, update, and remove notifications. Specify resourceVersion. [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: V1PodList [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] kwargs['_return_http_data_only'] = True [e2e-llm-inference-service] > return self.list_namespaced_pod_with_http_info(namespace, **kwargs) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/core_v1_api.py:15968: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] namespace = 'kserve-ci-e2e-test' [e2e-llm-inference-service] kwargs = {'_return_http_data_only': True, 'label_selector': 'app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload'} [e2e-llm-inference-service] local_var_params = {'_return_http_data_only': True, 'all_params': ['namespace', 'pretty', 'allow_watch_bookmarks', '_continue', 'field_selector', 'label_selector', ...], 'auth_settings': ['BearerToken'], 'body_params': None, ...} [e2e-llm-inference-service] all_params = ['namespace', 'pretty', 'allow_watch_bookmarks', '_continue', 'field_selector', 'label_selector', ...] [e2e-llm-inference-service] key = '_return_http_data_only', val = True, collection_formats = {} [e2e-llm-inference-service] path_params = {'namespace': 'kserve-ci-e2e-test'} [e2e-llm-inference-service] query_params = [('labelSelector', 'app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload')] [e2e-llm-inference-service] [e2e-llm-inference-service] def list_namespaced_pod_with_http_info(self, namespace, **kwargs): # noqa: E501 [e2e-llm-inference-service] """list_namespaced_pod # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] list or watch objects of kind Pod # noqa: E501 [e2e-llm-inference-service] This method makes a synchronous HTTP request by default. To make an [e2e-llm-inference-service] asynchronous HTTP request, please pass async_req=True [e2e-llm-inference-service] >>> thread = api.list_namespaced_pod_with_http_info(namespace, async_req=True) [e2e-llm-inference-service] >>> result = thread.get() [e2e-llm-inference-service] [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param str namespace: object name and auth scope, such as for teams and projects (required) [e2e-llm-inference-service] :param str pretty: If 'true', then the output is pretty printed. Defaults to 'false' unless the user-agent indicates a browser or command-line HTTP tool (curl and wget). [e2e-llm-inference-service] :param bool allow_watch_bookmarks: allowWatchBookmarks requests watch events with type \"BOOKMARK\". Servers that do not implement bookmarks may ignore this flag and bookmarks are sent at the server's discretion. Clients should not assume bookmarks are returned at any specific interval, nor may they assume the server will send any BOOKMARK event during a session. If this is not a watch, this field is ignored. [e2e-llm-inference-service] :param str _continue: The continue option should be set when retrieving more results from the server. Since this value is server defined, clients may only use the continue value from a previous query result with identical query parameters (except for the value of continue) and the server may reject a continue value it does not recognize. If the specified continue value is no longer valid whether due to expiration (generally five to fifteen minutes) or a configuration change on the server, the server will respond with a 410 ResourceExpired error together with a continue token. If the client needs a consistent list, it must restart their list without the continue field. Otherwise, the client may send another list request with the token received with the 410 error, the server will respond with a list starting from the next key, but from the latest snapshot, which is inconsistent from the previous list results - objects that are created, modified, or deleted after the first list request will be included in the response, as long as their keys are after the \"next key\". This field is not supported when watch is true. Clients may start a watch from the last resourceVersion value returned by the server and not miss any modifications. [e2e-llm-inference-service] :param str field_selector: A selector to restrict the list of returned objects by their fields. Defaults to everything. [e2e-llm-inference-service] :param str label_selector: A selector to restrict the list of returned objects by their labels. Defaults to everything. [e2e-llm-inference-service] :param int limit: limit is a maximum number of responses to return for a list call. If more items exist, the server will set the `continue` field on the list metadata to a value that can be used with the same initial query to retrieve the next set of results. Setting a limit may return fewer than the requested amount of items (up to zero items) in the event all requested objects are filtered out and clients should only use the presence of the continue field to determine whether more results are available. Servers may choose not to support the limit argument and will return all of the available results. If limit is specified and the continue field is empty, clients may assume that no more results are available. This field is not supported if watch is true. The server guarantees that the objects returned when using continue will be identical to issuing a single list call without a limit - that is, no objects created, modified, or deleted after the first request is issued will be included in any subsequent continued requests. This is sometimes referred to as a consistent snapshot, and ensures that a client that is using limit to receive smaller chunks of a very large result can ensure they see all possible objects. If objects are updated during a chunked list the version of the object that was present at the time the first list result was calculated is returned. [e2e-llm-inference-service] :param str resource_version: resourceVersion sets a constraint on what resource versions a request may be served from. See https://kubernetes.io/docs/reference/using-api/api-concepts/#resource-versions for details. Defaults to unset [e2e-llm-inference-service] :param str resource_version_match: resourceVersionMatch determines how resourceVersion is applied to list calls. It is highly recommended that resourceVersionMatch be set for list calls where resourceVersion is set See https://kubernetes.io/docs/reference/using-api/api-concepts/#resource-versions for details. Defaults to unset [e2e-llm-inference-service] :param bool send_initial_events: `sendInitialEvents=true` may be set together with `watch=true`. In that case, the watch stream will begin with synthetic events to produce the current state of objects in the collection. Once all such events have been sent, a synthetic \"Bookmark\" event will be sent. The bookmark will report the ResourceVersion (RV) corresponding to the set of objects, and be marked with `\"k8s.io/initial-events-end\": \"true\"` annotation. Afterwards, the watch stream will proceed as usual, sending watch events corresponding to changes (subsequent to the RV) to objects watched. When `sendInitialEvents` option is set, we require `resourceVersionMatch` option to also be set. The semantic of the watch request is as following: - `resourceVersionMatch` = NotOlderThan is interpreted as \"data at least as new as the provided `resourceVersion`\" and the bookmark event is send when the state is synced to a `resourceVersion` at least as fresh as the one provided by the ListOptions. If `resourceVersion` is unset, this is interpreted as \"consistent read\" and the bookmark event is send when the state is synced at least to the moment when request started being processed. - `resourceVersionMatch` set to any other value or unset Invalid error is returned. Defaults to true if `resourceVersion=\"\"` or `resourceVersion=\"0\"` (for backward compatibility reasons) and to false otherwise. [e2e-llm-inference-service] :param int timeout_seconds: Timeout for the list/watch call. This limits the duration of the call, regardless of any activity or inactivity. [e2e-llm-inference-service] :param bool watch: Watch for changes to the described resources and return them as a stream of add, update, and remove notifications. Specify resourceVersion. [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: tuple(V1PodList, status_code(int), headers(HTTPHeaderDict)) [e2e-llm-inference-service] If the method is called asynchronously, [e2e-llm-inference-service] returns the request thread. [e2e-llm-inference-service] """ [e2e-llm-inference-service] [e2e-llm-inference-service] local_var_params = locals() [e2e-llm-inference-service] [e2e-llm-inference-service] all_params = [ [e2e-llm-inference-service] 'namespace', [e2e-llm-inference-service] 'pretty', [e2e-llm-inference-service] 'allow_watch_bookmarks', [e2e-llm-inference-service] '_continue', [e2e-llm-inference-service] 'field_selector', [e2e-llm-inference-service] 'label_selector', [e2e-llm-inference-service] 'limit', [e2e-llm-inference-service] 'resource_version', [e2e-llm-inference-service] 'resource_version_match', [e2e-llm-inference-service] 'send_initial_events', [e2e-llm-inference-service] 'timeout_seconds', [e2e-llm-inference-service] 'watch' [e2e-llm-inference-service] ] [e2e-llm-inference-service] all_params.extend( [e2e-llm-inference-service] [ [e2e-llm-inference-service] 'async_req', [e2e-llm-inference-service] '_return_http_data_only', [e2e-llm-inference-service] '_preload_content', [e2e-llm-inference-service] '_request_timeout' [e2e-llm-inference-service] ] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] for key, val in six.iteritems(local_var_params['kwargs']): [e2e-llm-inference-service] if key not in all_params: [e2e-llm-inference-service] raise ApiTypeError( [e2e-llm-inference-service] "Got an unexpected keyword argument '%s'" [e2e-llm-inference-service] " to method list_namespaced_pod" % key [e2e-llm-inference-service] ) [e2e-llm-inference-service] local_var_params[key] = val [e2e-llm-inference-service] del local_var_params['kwargs'] [e2e-llm-inference-service] # verify the required parameter 'namespace' is set [e2e-llm-inference-service] if self.api_client.client_side_validation and ('namespace' not in local_var_params or # noqa: E501 [e2e-llm-inference-service] local_var_params['namespace'] is None): # noqa: E501 [e2e-llm-inference-service] raise ApiValueError("Missing the required parameter `namespace` when calling `list_namespaced_pod`") # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] collection_formats = {} [e2e-llm-inference-service] [e2e-llm-inference-service] path_params = {} [e2e-llm-inference-service] if 'namespace' in local_var_params: [e2e-llm-inference-service] path_params['namespace'] = local_var_params['namespace'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] query_params = [] [e2e-llm-inference-service] if 'pretty' in local_var_params and local_var_params['pretty'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('pretty', local_var_params['pretty'])) # noqa: E501 [e2e-llm-inference-service] if 'allow_watch_bookmarks' in local_var_params and local_var_params['allow_watch_bookmarks'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('allowWatchBookmarks', local_var_params['allow_watch_bookmarks'])) # noqa: E501 [e2e-llm-inference-service] if '_continue' in local_var_params and local_var_params['_continue'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('continue', local_var_params['_continue'])) # noqa: E501 [e2e-llm-inference-service] if 'field_selector' in local_var_params and local_var_params['field_selector'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('fieldSelector', local_var_params['field_selector'])) # noqa: E501 [e2e-llm-inference-service] if 'label_selector' in local_var_params and local_var_params['label_selector'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('labelSelector', local_var_params['label_selector'])) # noqa: E501 [e2e-llm-inference-service] if 'limit' in local_var_params and local_var_params['limit'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('limit', local_var_params['limit'])) # noqa: E501 [e2e-llm-inference-service] if 'resource_version' in local_var_params and local_var_params['resource_version'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('resourceVersion', local_var_params['resource_version'])) # noqa: E501 [e2e-llm-inference-service] if 'resource_version_match' in local_var_params and local_var_params['resource_version_match'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('resourceVersionMatch', local_var_params['resource_version_match'])) # noqa: E501 [e2e-llm-inference-service] if 'send_initial_events' in local_var_params and local_var_params['send_initial_events'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('sendInitialEvents', local_var_params['send_initial_events'])) # noqa: E501 [e2e-llm-inference-service] if 'timeout_seconds' in local_var_params and local_var_params['timeout_seconds'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('timeoutSeconds', local_var_params['timeout_seconds'])) # noqa: E501 [e2e-llm-inference-service] if 'watch' in local_var_params and local_var_params['watch'] is not None: # noqa: E501 [e2e-llm-inference-service] query_params.append(('watch', local_var_params['watch'])) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] header_params = {} [e2e-llm-inference-service] [e2e-llm-inference-service] form_params = [] [e2e-llm-inference-service] local_var_files = {} [e2e-llm-inference-service] [e2e-llm-inference-service] body_params = None [e2e-llm-inference-service] # HTTP header `Accept` [e2e-llm-inference-service] header_params['Accept'] = self.api_client.select_header_accept( [e2e-llm-inference-service] ['application/json', 'application/yaml', 'application/vnd.kubernetes.protobuf', 'application/cbor', 'application/json;stream=watch', 'application/vnd.kubernetes.protobuf;stream=watch', 'application/cbor-seq']) # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] # Authentication setting [e2e-llm-inference-service] auth_settings = ['BearerToken'] # noqa: E501 [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.api_client.call_api( [e2e-llm-inference-service] '/api/v1/namespaces/{namespace}/pods', 'GET', [e2e-llm-inference-service] path_params, [e2e-llm-inference-service] query_params, [e2e-llm-inference-service] header_params, [e2e-llm-inference-service] body=body_params, [e2e-llm-inference-service] post_params=form_params, [e2e-llm-inference-service] files=local_var_files, [e2e-llm-inference-service] response_type='V1PodList', # noqa: E501 [e2e-llm-inference-service] auth_settings=auth_settings, [e2e-llm-inference-service] async_req=local_var_params.get('async_req'), [e2e-llm-inference-service] _return_http_data_only=local_var_params.get('_return_http_data_only'), # noqa: E501 [e2e-llm-inference-service] _preload_content=local_var_params.get('_preload_content', True), [e2e-llm-inference-service] _request_timeout=local_var_params.get('_request_timeout'), [e2e-llm-inference-service] collection_formats=collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api/core_v1_api.py:16087: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/api/v1/namespaces/{namespace}/pods', method = 'GET' [e2e-llm-inference-service] path_params = {'namespace': 'kserve-ci-e2e-test'} [e2e-llm-inference-service] query_params = [('labelSelector', 'app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload')] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'V1PodList' [e2e-llm-inference-service] auth_settings = ['BearerToken'], async_req = None, _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def call_api(self, resource_path, method, [e2e-llm-inference-service] path_params=None, query_params=None, header_params=None, [e2e-llm-inference-service] body=None, post_params=None, files=None, [e2e-llm-inference-service] response_type=None, auth_settings=None, async_req=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] """Makes the HTTP request (synchronous) and returns deserialized data. [e2e-llm-inference-service] [e2e-llm-inference-service] To make an async_req request, set the async_req parameter. [e2e-llm-inference-service] [e2e-llm-inference-service] :param resource_path: Path to method endpoint. [e2e-llm-inference-service] :param method: Method to call. [e2e-llm-inference-service] :param path_params: Path parameters in the url. [e2e-llm-inference-service] :param query_params: Query parameters in the url. [e2e-llm-inference-service] :param header_params: Header parameters to be [e2e-llm-inference-service] placed in the request header. [e2e-llm-inference-service] :param body: Request body. [e2e-llm-inference-service] :param post_params dict: Request post form parameters, [e2e-llm-inference-service] for `application/x-www-form-urlencoded`, `multipart/form-data`. [e2e-llm-inference-service] :param auth_settings list: Auth Settings names for the request. [e2e-llm-inference-service] :param response: Response data type. [e2e-llm-inference-service] :param files dict: key -> filename, value -> filepath, [e2e-llm-inference-service] for `multipart/form-data`. [e2e-llm-inference-service] :param async_req bool: execute request asynchronously [e2e-llm-inference-service] :param _return_http_data_only: response data without head status code [e2e-llm-inference-service] and headers [e2e-llm-inference-service] :param collection_formats: dict of collection formats for path, query, [e2e-llm-inference-service] header, and post parameters. [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] :return: [e2e-llm-inference-service] If async_req parameter is True, [e2e-llm-inference-service] the request will be called asynchronously. [e2e-llm-inference-service] The method will return the request thread. [e2e-llm-inference-service] If parameter async_req is False or missing, [e2e-llm-inference-service] then the method will return the response directly. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if not async_req: [e2e-llm-inference-service] > return self.__call_api(resource_path, method, [e2e-llm-inference-service] path_params, query_params, header_params, [e2e-llm-inference-service] body, post_params, files, [e2e-llm-inference-service] response_type, auth_settings, [e2e-llm-inference-service] _return_http_data_only, collection_formats, [e2e-llm-inference-service] _preload_content, _request_timeout, _host) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:348: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] resource_path = '/api/v1/namespaces/kserve-ci-e2e-test/pods', method = 'GET' [e2e-llm-inference-service] path_params = [('namespace', 'kserve-ci-e2e-test')] [e2e-llm-inference-service] query_params = [('labelSelector', 'app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload')] [e2e-llm-inference-service] header_params = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = [], files = {}, response_type = 'V1PodList' [e2e-llm-inference-service] auth_settings = ['BearerToken'], _return_http_data_only = True [e2e-llm-inference-service] collection_formats = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] _host = None [e2e-llm-inference-service] [e2e-llm-inference-service] def __call_api( [e2e-llm-inference-service] self, resource_path, method, path_params=None, [e2e-llm-inference-service] query_params=None, header_params=None, body=None, post_params=None, [e2e-llm-inference-service] files=None, response_type=None, auth_settings=None, [e2e-llm-inference-service] _return_http_data_only=None, collection_formats=None, [e2e-llm-inference-service] _preload_content=True, _request_timeout=None, _host=None): [e2e-llm-inference-service] [e2e-llm-inference-service] config = self.configuration [e2e-llm-inference-service] [e2e-llm-inference-service] # header parameters [e2e-llm-inference-service] header_params = header_params or {} [e2e-llm-inference-service] header_params.update(self.default_headers) [e2e-llm-inference-service] if self.cookie: [e2e-llm-inference-service] header_params['Cookie'] = self.cookie [e2e-llm-inference-service] if header_params: [e2e-llm-inference-service] header_params = self.sanitize_for_serialization(header_params) [e2e-llm-inference-service] header_params = dict(self.parameters_to_tuples(header_params, [e2e-llm-inference-service] collection_formats)) [e2e-llm-inference-service] [e2e-llm-inference-service] # path parameters [e2e-llm-inference-service] if path_params: [e2e-llm-inference-service] path_params = self.sanitize_for_serialization(path_params) [e2e-llm-inference-service] path_params = self.parameters_to_tuples(path_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] for k, v in path_params: [e2e-llm-inference-service] # specified safe chars, encode everything [e2e-llm-inference-service] resource_path = resource_path.replace( [e2e-llm-inference-service] '{%s}' % k, [e2e-llm-inference-service] quote(str(v), safe=config.safe_chars_for_path_param) [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # query parameters [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] query_params = self.sanitize_for_serialization(query_params) [e2e-llm-inference-service] query_params = self.parameters_to_tuples(query_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] [e2e-llm-inference-service] # post parameters [e2e-llm-inference-service] if post_params or files: [e2e-llm-inference-service] post_params = post_params if post_params else [] [e2e-llm-inference-service] post_params = self.sanitize_for_serialization(post_params) [e2e-llm-inference-service] post_params = self.parameters_to_tuples(post_params, [e2e-llm-inference-service] collection_formats) [e2e-llm-inference-service] post_params.extend(self.files_parameters(files)) [e2e-llm-inference-service] [e2e-llm-inference-service] # auth setting [e2e-llm-inference-service] self.update_params_for_auth(header_params, query_params, auth_settings) [e2e-llm-inference-service] [e2e-llm-inference-service] # body [e2e-llm-inference-service] if body: [e2e-llm-inference-service] body = self.sanitize_for_serialization(body) [e2e-llm-inference-service] [e2e-llm-inference-service] # request url [e2e-llm-inference-service] if _host is None: [e2e-llm-inference-service] url = self.configuration.host + resource_path [e2e-llm-inference-service] else: [e2e-llm-inference-service] # use server/host defined in path or operation instead [e2e-llm-inference-service] url = _host + resource_path [e2e-llm-inference-service] [e2e-llm-inference-service] # perform request and return response [e2e-llm-inference-service] > response_data = self.request( [e2e-llm-inference-service] method, url, query_params=query_params, headers=header_params, [e2e-llm-inference-service] post_params=post_params, body=body, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:180: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api/v1/namespaces/kserve-ci-e2e-test/pods' [e2e-llm-inference-service] query_params = [('labelSelector', 'app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload')] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] post_params = [], body = None, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] post_params=None, body=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Makes the HTTP request using RESTClient.""" [e2e-llm-inference-service] if method == "GET": [e2e-llm-inference-service] > return self.rest_client.GET(url, [e2e-llm-inference-service] query_params=query_params, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/api_client.py:373: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api/v1/namespaces/kserve-ci-e2e-test/pods' [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] query_params = [('labelSelector', 'app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload')] [e2e-llm-inference-service] _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def GET(self, url, headers=None, query_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] > return self.request("GET", url, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] _preload_content=_preload_content, [e2e-llm-inference-service] _request_timeout=_request_timeout, [e2e-llm-inference-service] query_params=query_params) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:244: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api/v1/namespaces/kserve-ci-e2e-test/pods' [e2e-llm-inference-service] query_params = [('labelSelector', 'app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload')] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] body = None, post_params = {}, _preload_content = True, _request_timeout = None [e2e-llm-inference-service] [e2e-llm-inference-service] def request(self, method, url, query_params=None, headers=None, [e2e-llm-inference-service] body=None, post_params=None, _preload_content=True, [e2e-llm-inference-service] _request_timeout=None): [e2e-llm-inference-service] """Perform requests. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: http request method [e2e-llm-inference-service] :param url: http request url [e2e-llm-inference-service] :param query_params: query parameters in the url [e2e-llm-inference-service] :param headers: http request headers [e2e-llm-inference-service] :param body: request json body, for `application/json` [e2e-llm-inference-service] :param post_params: request post parameters, [e2e-llm-inference-service] `application/x-www-form-urlencoded` [e2e-llm-inference-service] and `multipart/form-data` [e2e-llm-inference-service] :param _preload_content: if False, the urllib3.HTTPResponse object will [e2e-llm-inference-service] be returned without reading/decoding response [e2e-llm-inference-service] data. Default is True. [e2e-llm-inference-service] :param _request_timeout: timeout setting for this request. If one [e2e-llm-inference-service] number provided, it will be total request [e2e-llm-inference-service] timeout. It can also be a pair (tuple) of [e2e-llm-inference-service] (connection, read) timeouts. [e2e-llm-inference-service] """ [e2e-llm-inference-service] method = method.upper() [e2e-llm-inference-service] assert method in ['GET', 'HEAD', 'DELETE', 'POST', 'PUT', [e2e-llm-inference-service] 'PATCH', 'OPTIONS'] [e2e-llm-inference-service] [e2e-llm-inference-service] if post_params and body: [e2e-llm-inference-service] raise ApiValueError( [e2e-llm-inference-service] "body parameter cannot be used with post_params parameter." [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] post_params = post_params or {} [e2e-llm-inference-service] headers = headers or {} [e2e-llm-inference-service] [e2e-llm-inference-service] timeout = None [e2e-llm-inference-service] if _request_timeout: [e2e-llm-inference-service] if isinstance(_request_timeout, (int, ) if six.PY3 else (int, long)): # noqa: E501,F821 [e2e-llm-inference-service] timeout = urllib3.Timeout(total=_request_timeout) [e2e-llm-inference-service] elif (isinstance(_request_timeout, tuple) and [e2e-llm-inference-service] len(_request_timeout) == 2): [e2e-llm-inference-service] timeout = urllib3.Timeout( [e2e-llm-inference-service] connect=_request_timeout[0], read=_request_timeout[1]) [e2e-llm-inference-service] [e2e-llm-inference-service] if 'Content-Type' not in headers: [e2e-llm-inference-service] headers['Content-Type'] = 'application/json' [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # For `POST`, `PUT`, `PATCH`, `OPTIONS`, `DELETE` [e2e-llm-inference-service] if method in ['POST', 'PUT', 'PATCH', 'OPTIONS', 'DELETE']: [e2e-llm-inference-service] if query_params: [e2e-llm-inference-service] url += '?' + urlencode(query_params) [e2e-llm-inference-service] if (re.search('json', headers['Content-Type'], re.IGNORECASE) or [e2e-llm-inference-service] headers['Content-Type'] == 'application/apply-patch+yaml'): [e2e-llm-inference-service] if headers['Content-Type'] == 'application/json-patch+json': [e2e-llm-inference-service] if not isinstance(body, list): [e2e-llm-inference-service] headers['Content-Type'] = \ [e2e-llm-inference-service] 'application/strategic-merge-patch+json' [e2e-llm-inference-service] request_body = None [e2e-llm-inference-service] if body is not None: [e2e-llm-inference-service] request_body = json.dumps(body) [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'application/x-www-form-urlencoded': # noqa: E501 [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=False, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] elif headers['Content-Type'] == 'multipart/form-data': [e2e-llm-inference-service] # must del headers['Content-Type'], or the correct [e2e-llm-inference-service] # Content-Type which generated by urllib3 will be [e2e-llm-inference-service] # overwritten. [e2e-llm-inference-service] del headers['Content-Type'] [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] fields=post_params, [e2e-llm-inference-service] encode_multipart=True, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] # Pass a `string` parameter directly in the body to support [e2e-llm-inference-service] # other content types than Json when `body` argument is [e2e-llm-inference-service] # provided in serialized form [e2e-llm-inference-service] elif isinstance(body, str) or isinstance(body, bytes): [e2e-llm-inference-service] request_body = body [e2e-llm-inference-service] r = self.pool_manager.request( [e2e-llm-inference-service] method, url, [e2e-llm-inference-service] body=request_body, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Cannot generate the request from given parameters [e2e-llm-inference-service] msg = """Cannot prepare a request message for provided [e2e-llm-inference-service] arguments. Please check that your arguments match [e2e-llm-inference-service] declared content type.""" [e2e-llm-inference-service] raise ApiException(status=0, reason=msg) [e2e-llm-inference-service] # For `GET`, `HEAD` [e2e-llm-inference-service] else: [e2e-llm-inference-service] > r = self.pool_manager.request(method, url, [e2e-llm-inference-service] fields=query_params, [e2e-llm-inference-service] preload_content=_preload_content, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] headers=headers) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/kubernetes/client/rest.py:217: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api/v1/namespaces/kserve-ci-e2e-test/pods' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] fields = [('labelSelector', 'app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload')] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] json = None, urlopen_kw = {'preload_content': True, 'timeout': None} [e2e-llm-inference-service] [e2e-llm-inference-service] def request( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] fields: _TYPE_FIELDS | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] json: typing.Any | None = None, [e2e-llm-inference-service] **urlopen_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Make a request using :meth:`urlopen` with the appropriate encoding of [e2e-llm-inference-service] ``fields`` based on the ``method`` used. [e2e-llm-inference-service] [e2e-llm-inference-service] This is a convenience method that requires the least amount of manual [e2e-llm-inference-service] effort. It can be used in most situations, while still having the [e2e-llm-inference-service] option to drop down to more specific methods when necessary, such as [e2e-llm-inference-service] :meth:`request_encode_url`, :meth:`request_encode_body`, [e2e-llm-inference-service] or even the lowest level :meth:`urlopen`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param fields: [e2e-llm-inference-service] Data to encode and send in the URL or request body, depending on ``method``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param json: [e2e-llm-inference-service] Data to encode and send as JSON with UTF-encoded in the request body. [e2e-llm-inference-service] The ``"Content-Type"`` header will be set to ``"application/json"`` [e2e-llm-inference-service] unless specified otherwise. [e2e-llm-inference-service] """ [e2e-llm-inference-service] method = method.upper() [e2e-llm-inference-service] [e2e-llm-inference-service] if json is not None and body is not None: [e2e-llm-inference-service] raise TypeError( [e2e-llm-inference-service] "request got values for both 'body' and 'json' parameters which are mutually exclusive" [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if json is not None: [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not ("content-type" in map(str.lower, headers.keys())): [e2e-llm-inference-service] headers = HTTPHeaderDict(headers) [e2e-llm-inference-service] headers["Content-Type"] = "application/json" [e2e-llm-inference-service] [e2e-llm-inference-service] body = _json.dumps(json, separators=(",", ":"), ensure_ascii=False).encode( [e2e-llm-inference-service] "utf-8" [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if body is not None: [e2e-llm-inference-service] urlopen_kw["body"] = body [e2e-llm-inference-service] [e2e-llm-inference-service] if method in self._encode_url_methods: [e2e-llm-inference-service] > return self.request_encode_url( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] fields=fields, # type: ignore[arg-type] [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] **urlopen_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/_request_methods.py:135: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload' [e2e-llm-inference-service] fields = [('labelSelector', 'app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload')] [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] urlopen_kw = {'preload_content': True, 'timeout': None} [e2e-llm-inference-service] extra_kw = {'headers': {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'}, 'preload_content': True, 'timeout': None} [e2e-llm-inference-service] [e2e-llm-inference-service] def request_encode_url( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] fields: _TYPE_ENCODE_URL_FIELDS | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] **urlopen_kw: str, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Make a request using :meth:`urlopen` with the ``fields`` encoded in [e2e-llm-inference-service] the url. This is useful for request methods like GET, HEAD, DELETE, etc. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param fields: [e2e-llm-inference-service] Data to encode and send in the URL. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] extra_kw: dict[str, typing.Any] = {"headers": headers} [e2e-llm-inference-service] extra_kw.update(urlopen_kw) [e2e-llm-inference-service] [e2e-llm-inference-service] if fields: [e2e-llm-inference-service] url += "?" + urlencode(fields) [e2e-llm-inference-service] [e2e-llm-inference-service] > return self.urlopen(method, url, **extra_kw) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/_request_methods.py:182: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = 'https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload' [e2e-llm-inference-service] redirect = True [e2e-llm-inference-service] kw = {'assert_same_host': False, 'headers': {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'}, 'preload_content': True, 'redirect': False, ...} [e2e-llm-inference-service] u = Url(scheme='https', auth=None, host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', p..., query='labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload', fragment=None) [e2e-llm-inference-service] conn = [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, method: str, url: str, redirect: bool = True, **kw: typing.Any [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Same as :meth:`urllib3.HTTPConnectionPool.urlopen` [e2e-llm-inference-service] with custom cross-host redirect logic and only sends the request-uri [e2e-llm-inference-service] portion of the ``url``. [e2e-llm-inference-service] [e2e-llm-inference-service] The given ``url`` parameter must be absolute, such that an appropriate [e2e-llm-inference-service] :class:`urllib3.connectionpool.ConnectionPool` can be chosen for it. [e2e-llm-inference-service] """ [e2e-llm-inference-service] u = parse_url(url) [e2e-llm-inference-service] [e2e-llm-inference-service] if u.scheme is None: [e2e-llm-inference-service] warnings.warn( [e2e-llm-inference-service] "URLs without a scheme (ie 'https://') are deprecated and will raise an error " [e2e-llm-inference-service] "in a future version of urllib3. To avoid this DeprecationWarning ensure all URLs " [e2e-llm-inference-service] "start with 'https://' or 'http://'. Read more in this issue: " [e2e-llm-inference-service] "https://github.com/urllib3/urllib3/issues/2920", [e2e-llm-inference-service] category=DeprecationWarning, [e2e-llm-inference-service] stacklevel=2, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = self.connection_from_host(u.host, port=u.port, scheme=u.scheme) [e2e-llm-inference-service] [e2e-llm-inference-service] kw["assert_same_host"] = False [e2e-llm-inference-service] kw["redirect"] = False [e2e-llm-inference-service] [e2e-llm-inference-service] if "headers" not in kw: [e2e-llm-inference-service] kw["headers"] = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if self._proxy_requires_url_absolute_form(u): [e2e-llm-inference-service] response = conn.urlopen(method, url, **kw) [e2e-llm-inference-service] else: [e2e-llm-inference-service] > response = conn.urlopen(method, u.request_uri, **kw) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/poolmanager.py:457: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=2, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False, timeout = None, pool_timeout = None [e2e-llm-inference-service] release_conn = True, chunked = False, body_pos = None, preload_content = True [e2e-llm-inference-service] decode_content = True, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/api/v1/namespaces/kserve-ci-e2e-test/pods', query='labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload', fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ConnectionResetError(104, 'Connection reset by peer'), clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=1, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False, timeout = None, pool_timeout = None [e2e-llm-inference-service] release_conn = True, chunked = False, body_pos = None, preload_content = True [e2e-llm-inference-service] decode_content = True, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/api/v1/namespaces/kserve-ci-e2e-test/pods', query='labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload', fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = ConnectTimeoutError( BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False, timeout = None, pool_timeout = None [e2e-llm-inference-service] release_conn = True, chunked = False, body_pos = None, preload_content = True [e2e-llm-inference-service] decode_content = True, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/api/v1/namespaces/kserve-ci-e2e-test/pods', query='labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload', fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False [e2e-llm-inference-service] err = NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.c...a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)") [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] retries.sleep() [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of the error for the retry warning. [e2e-llm-inference-service] err = e [e2e-llm-inference-service] [e2e-llm-inference-service] finally: [e2e-llm-inference-service] if not clean_exit: [e2e-llm-inference-service] # We hit some kind of exception, handled or otherwise. We need [e2e-llm-inference-service] # to throw the connection away unless explicitly told not to. [e2e-llm-inference-service] # Close the connection, set the variable to None, and make sure [e2e-llm-inference-service] # we put the None back in the pool to avoid leaking it. [e2e-llm-inference-service] if conn: [e2e-llm-inference-service] conn.close() [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] release_this_conn = True [e2e-llm-inference-service] [e2e-llm-inference-service] if release_this_conn: [e2e-llm-inference-service] # Put the connection back to be reused. If the connection is [e2e-llm-inference-service] # expired then it will be None, which will get replaced with a [e2e-llm-inference-service] # fresh connection during _get_conn. [e2e-llm-inference-service] self._put_conn(conn) [e2e-llm-inference-service] [e2e-llm-inference-service] if not conn: [e2e-llm-inference-service] # Try again [e2e-llm-inference-service] log.warning( [e2e-llm-inference-service] "Retrying (%r) after connection broken by '%r': %s", retries, err, url [e2e-llm-inference-service] ) [e2e-llm-inference-service] > return self.urlopen( [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] body, [e2e-llm-inference-service] headers, [e2e-llm-inference-service] retries, [e2e-llm-inference-service] redirect, [e2e-llm-inference-service] assert_same_host, [e2e-llm-inference-service] timeout=timeout, [e2e-llm-inference-service] pool_timeout=pool_timeout, [e2e-llm-inference-service] release_conn=release_conn, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] body_pos=body_pos, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:871: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload' [e2e-llm-inference-service] body = None [e2e-llm-inference-service] headers = {'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'OpenAPI-Generator/32.0.1/python'} [e2e-llm-inference-service] retries = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] redirect = False, assert_same_host = False, timeout = None, pool_timeout = None [e2e-llm-inference-service] release_conn = True, chunked = False, body_pos = None, preload_content = True [e2e-llm-inference-service] decode_content = True, response_kw = {} [e2e-llm-inference-service] parsed_url = Url(scheme=None, auth=None, host=None, port=None, path='/api/v1/namespaces/kserve-ci-e2e-test/pods', query='labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload', fragment=None) [e2e-llm-inference-service] destination_scheme = None, conn = None, release_this_conn = True [e2e-llm-inference-service] http_tunnel_required = False, err = None, clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] def urlopen( # type: ignore[override] [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str, [e2e-llm-inference-service] url: str, [e2e-llm-inference-service] body: _TYPE_BODY | None = None, [e2e-llm-inference-service] headers: typing.Mapping[str, str] | None = None, [e2e-llm-inference-service] retries: Retry | bool | int | None = None, [e2e-llm-inference-service] redirect: bool = True, [e2e-llm-inference-service] assert_same_host: bool = True, [e2e-llm-inference-service] timeout: _TYPE_TIMEOUT = _DEFAULT_TIMEOUT, [e2e-llm-inference-service] pool_timeout: int | None = None, [e2e-llm-inference-service] release_conn: bool | None = None, [e2e-llm-inference-service] chunked: bool = False, [e2e-llm-inference-service] body_pos: _TYPE_BODY_POSITION | None = None, [e2e-llm-inference-service] preload_content: bool = True, [e2e-llm-inference-service] decode_content: bool = True, [e2e-llm-inference-service] **response_kw: typing.Any, [e2e-llm-inference-service] ) -> BaseHTTPResponse: [e2e-llm-inference-service] """ [e2e-llm-inference-service] Get a connection from the pool and perform an HTTP request. This is the [e2e-llm-inference-service] lowest level call for making a request, so you'll need to specify all [e2e-llm-inference-service] the raw details. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] More commonly, it's appropriate to use a convenience method [e2e-llm-inference-service] such as :meth:`request`. [e2e-llm-inference-service] [e2e-llm-inference-service] .. note:: [e2e-llm-inference-service] [e2e-llm-inference-service] `release_conn` will only behave as expected if [e2e-llm-inference-service] `preload_content=False` because we want to make [e2e-llm-inference-service] `preload_content=False` the default behaviour someday soon without [e2e-llm-inference-service] breaking backwards compatibility. [e2e-llm-inference-service] [e2e-llm-inference-service] :param method: [e2e-llm-inference-service] HTTP request method (such as GET, POST, PUT, etc.) [e2e-llm-inference-service] [e2e-llm-inference-service] :param url: [e2e-llm-inference-service] The URL to perform the request on. [e2e-llm-inference-service] [e2e-llm-inference-service] :param body: [e2e-llm-inference-service] Data to send in the request body, either :class:`str`, :class:`bytes`, [e2e-llm-inference-service] an iterable of :class:`str`/:class:`bytes`, or a file-like object. [e2e-llm-inference-service] [e2e-llm-inference-service] :param headers: [e2e-llm-inference-service] Dictionary of custom headers to send, such as User-Agent, [e2e-llm-inference-service] If-None-Match, etc. If None, pool headers are used. If provided, [e2e-llm-inference-service] these headers completely replace any pool-specific headers. [e2e-llm-inference-service] [e2e-llm-inference-service] :param retries: [e2e-llm-inference-service] Configure the number of retries to allow before raising a [e2e-llm-inference-service] :class:`~urllib3.exceptions.MaxRetryError` exception. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``None`` (default) will retry 3 times, see ``Retry.DEFAULT``. Pass a [e2e-llm-inference-service] :class:`~urllib3.util.retry.Retry` object for fine-grained control [e2e-llm-inference-service] over different types of retries. [e2e-llm-inference-service] Pass an integer number to retry connection errors that many times, [e2e-llm-inference-service] but no other types of errors. Pass zero to never retry. [e2e-llm-inference-service] [e2e-llm-inference-service] If ``False``, then retries are disabled and any exception is raised [e2e-llm-inference-service] immediately. Also, instead of raising a MaxRetryError on redirects, [e2e-llm-inference-service] the redirect response will be returned. [e2e-llm-inference-service] [e2e-llm-inference-service] :type retries: :class:`~urllib3.util.retry.Retry`, False, or an int. [e2e-llm-inference-service] [e2e-llm-inference-service] :param redirect: [e2e-llm-inference-service] If True, automatically handle redirects (status codes 301, 302, [e2e-llm-inference-service] 303, 307, 308). Each redirect counts as a retry. Disabling retries [e2e-llm-inference-service] will disable redirect, too. [e2e-llm-inference-service] [e2e-llm-inference-service] :param assert_same_host: [e2e-llm-inference-service] If ``True``, will make sure that the host of the pool requests is [e2e-llm-inference-service] consistent else will raise HostChangedError. When ``False``, you can [e2e-llm-inference-service] use the pool on an HTTP proxy and request foreign hosts. [e2e-llm-inference-service] [e2e-llm-inference-service] :param timeout: [e2e-llm-inference-service] If specified, overrides the default timeout for this one [e2e-llm-inference-service] request. It may be a float (in seconds) or an instance of [e2e-llm-inference-service] :class:`urllib3.util.Timeout`. [e2e-llm-inference-service] [e2e-llm-inference-service] :param pool_timeout: [e2e-llm-inference-service] If set and the pool is set to block=True, then this method will [e2e-llm-inference-service] block for ``pool_timeout`` seconds and raise EmptyPoolError if no [e2e-llm-inference-service] connection is available within the time period. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool preload_content: [e2e-llm-inference-service] If True, the response's body will be preloaded into memory. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool decode_content: [e2e-llm-inference-service] If True, will attempt to decode the body based on the [e2e-llm-inference-service] 'content-encoding' header. [e2e-llm-inference-service] [e2e-llm-inference-service] :param release_conn: [e2e-llm-inference-service] If False, then the urlopen call will not release the connection [e2e-llm-inference-service] back into the pool once a response is received (but will release if [e2e-llm-inference-service] you read the entire contents of the response such as when [e2e-llm-inference-service] `preload_content=True`). This is useful if you're not preloading [e2e-llm-inference-service] the response's content immediately. You will need to call [e2e-llm-inference-service] ``r.release_conn()`` on the response ``r`` to return the connection [e2e-llm-inference-service] back into the pool. If None, it takes the value of ``preload_content`` [e2e-llm-inference-service] which defaults to ``True``. [e2e-llm-inference-service] [e2e-llm-inference-service] :param bool chunked: [e2e-llm-inference-service] If True, urllib3 will send the body using chunked transfer [e2e-llm-inference-service] encoding. Otherwise, urllib3 will send the body using the standard [e2e-llm-inference-service] content-length form. Defaults to False. [e2e-llm-inference-service] [e2e-llm-inference-service] :param int body_pos: [e2e-llm-inference-service] Position to seek to in file-like body in the event of a retry or [e2e-llm-inference-service] redirect. Typically this won't need to be set because urllib3 will [e2e-llm-inference-service] auto-populate the value when needed. [e2e-llm-inference-service] """ [e2e-llm-inference-service] parsed_url = parse_url(url) [e2e-llm-inference-service] destination_scheme = parsed_url.scheme [e2e-llm-inference-service] [e2e-llm-inference-service] if headers is None: [e2e-llm-inference-service] headers = self.headers [e2e-llm-inference-service] [e2e-llm-inference-service] if not isinstance(retries, Retry): [e2e-llm-inference-service] retries = Retry.from_int(retries, redirect=redirect, default=self.retries) [e2e-llm-inference-service] [e2e-llm-inference-service] if release_conn is None: [e2e-llm-inference-service] release_conn = preload_content [e2e-llm-inference-service] [e2e-llm-inference-service] # Check host [e2e-llm-inference-service] if assert_same_host and not self.is_same_host(url): [e2e-llm-inference-service] raise HostChangedError(self, url, retries) [e2e-llm-inference-service] [e2e-llm-inference-service] # Ensure that the URL we're connecting to is properly encoded [e2e-llm-inference-service] if url.startswith("/"): [e2e-llm-inference-service] url = to_str(_encode_target(url)) [e2e-llm-inference-service] else: [e2e-llm-inference-service] url = to_str(parsed_url.url) [e2e-llm-inference-service] [e2e-llm-inference-service] conn = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Track whether `conn` needs to be released before [e2e-llm-inference-service] # returning/raising/recursing. Update this variable if necessary, and [e2e-llm-inference-service] # leave `release_conn` constant throughout the function. That way, if [e2e-llm-inference-service] # the function recurses, the original value of `release_conn` will be [e2e-llm-inference-service] # passed down into the recursive call, and its value will be respected. [e2e-llm-inference-service] # [e2e-llm-inference-service] # See issue #651 [1] for details. [e2e-llm-inference-service] # [e2e-llm-inference-service] # [1] [e2e-llm-inference-service] release_this_conn = release_conn [e2e-llm-inference-service] [e2e-llm-inference-service] http_tunnel_required = connection_requires_http_tunnel( [e2e-llm-inference-service] self.proxy, self.proxy_config, destination_scheme [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Merge the proxy headers. Only done when not using HTTP CONNECT. We [e2e-llm-inference-service] # have to copy the headers dict so we can safely change it without those [e2e-llm-inference-service] # changes being reflected in anyone else's copy. [e2e-llm-inference-service] if not http_tunnel_required: [e2e-llm-inference-service] headers = headers.copy() # type: ignore[attr-defined] [e2e-llm-inference-service] headers.update(self.proxy_headers) # type: ignore[union-attr] [e2e-llm-inference-service] [e2e-llm-inference-service] # Must keep the exception bound to a separate variable or else Python 3 [e2e-llm-inference-service] # complains about UnboundLocalError. [e2e-llm-inference-service] err = None [e2e-llm-inference-service] [e2e-llm-inference-service] # Keep track of whether we cleanly exited the except block. This [e2e-llm-inference-service] # ensures we do proper cleanup in finally. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] [e2e-llm-inference-service] # Rewind body position, if needed. Record current position [e2e-llm-inference-service] # for future rewinds in the event of a redirect/retry. [e2e-llm-inference-service] body_pos = set_file_position(body, body_pos) [e2e-llm-inference-service] [e2e-llm-inference-service] try: [e2e-llm-inference-service] # Request a connection from the queue. [e2e-llm-inference-service] timeout_obj = self._get_timeout(timeout) [e2e-llm-inference-service] conn = self._get_conn(timeout=pool_timeout) [e2e-llm-inference-service] [e2e-llm-inference-service] conn.timeout = timeout_obj.connect_timeout # type: ignore[assignment] [e2e-llm-inference-service] [e2e-llm-inference-service] # Is this a closed/new connection that requires CONNECT tunnelling? [e2e-llm-inference-service] if self.proxy is not None and http_tunnel_required and conn.is_closed: [e2e-llm-inference-service] try: [e2e-llm-inference-service] self._prepare_proxy(conn) [e2e-llm-inference-service] except (BaseSSLError, OSError, SocketTimeout) as e: [e2e-llm-inference-service] self._raise_timeout( [e2e-llm-inference-service] err=e, url=self.proxy.url, timeout_value=conn.timeout [e2e-llm-inference-service] ) [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] # If we're going to release the connection in ``finally:``, then [e2e-llm-inference-service] # the response doesn't need to know about the connection. Otherwise [e2e-llm-inference-service] # it will also try to release it and we'll have a double-release [e2e-llm-inference-service] # mess. [e2e-llm-inference-service] response_conn = conn if not release_conn else None [e2e-llm-inference-service] [e2e-llm-inference-service] # Make the request on the HTTPConnection object [e2e-llm-inference-service] response = self._make_request( [e2e-llm-inference-service] conn, [e2e-llm-inference-service] method, [e2e-llm-inference-service] url, [e2e-llm-inference-service] timeout=timeout_obj, [e2e-llm-inference-service] body=body, [e2e-llm-inference-service] headers=headers, [e2e-llm-inference-service] chunked=chunked, [e2e-llm-inference-service] retries=retries, [e2e-llm-inference-service] response_conn=response_conn, [e2e-llm-inference-service] preload_content=preload_content, [e2e-llm-inference-service] decode_content=decode_content, [e2e-llm-inference-service] **response_kw, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] # Everything went great! [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] [e2e-llm-inference-service] except EmptyPoolError: [e2e-llm-inference-service] # Didn't get a connection from the pool, no need to clean up [e2e-llm-inference-service] clean_exit = True [e2e-llm-inference-service] release_this_conn = False [e2e-llm-inference-service] raise [e2e-llm-inference-service] [e2e-llm-inference-service] except ( [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] ProtocolError, [e2e-llm-inference-service] BaseSSLError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] CertificateError, [e2e-llm-inference-service] ProxyError, [e2e-llm-inference-service] ) as e: [e2e-llm-inference-service] # Discard the connection for these exceptions. It will be [e2e-llm-inference-service] # replaced during the next _get_conn() call. [e2e-llm-inference-service] clean_exit = False [e2e-llm-inference-service] new_e: Exception = e [e2e-llm-inference-service] if isinstance(e, (BaseSSLError, CertificateError)): [e2e-llm-inference-service] new_e = SSLError(e) [e2e-llm-inference-service] if isinstance( [e2e-llm-inference-service] new_e, [e2e-llm-inference-service] ( [e2e-llm-inference-service] OSError, [e2e-llm-inference-service] NewConnectionError, [e2e-llm-inference-service] TimeoutError, [e2e-llm-inference-service] SSLError, [e2e-llm-inference-service] HTTPException, [e2e-llm-inference-service] ), [e2e-llm-inference-service] ) and (conn and conn.proxy and not conn.has_connected_to_proxy): [e2e-llm-inference-service] new_e = _wrap_proxy_error(new_e, conn.proxy.scheme) [e2e-llm-inference-service] elif isinstance(new_e, (OSError, HTTPException)): [e2e-llm-inference-service] new_e = ProtocolError("Connection aborted.", new_e) [e2e-llm-inference-service] [e2e-llm-inference-service] > retries = retries.increment( [e2e-llm-inference-service] method, url, error=new_e, _pool=self, _stacktrace=sys.exc_info()[2] [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/connectionpool.py:841: [e2e-llm-inference-service] _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ _ [e2e-llm-inference-service] [e2e-llm-inference-service] self = Retry(total=0, connect=None, read=None, redirect=None, status=None) [e2e-llm-inference-service] method = 'GET' [e2e-llm-inference-service] url = '/api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload' [e2e-llm-inference-service] response = None [e2e-llm-inference-service] error = NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.c...a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)") [e2e-llm-inference-service] _pool = [e2e-llm-inference-service] _stacktrace = [e2e-llm-inference-service] [e2e-llm-inference-service] def increment( [e2e-llm-inference-service] self, [e2e-llm-inference-service] method: str | None = None, [e2e-llm-inference-service] url: str | None = None, [e2e-llm-inference-service] response: BaseHTTPResponse | None = None, [e2e-llm-inference-service] error: Exception | None = None, [e2e-llm-inference-service] _pool: ConnectionPool | None = None, [e2e-llm-inference-service] _stacktrace: TracebackType | None = None, [e2e-llm-inference-service] ) -> Self: [e2e-llm-inference-service] """Return a new Retry object with incremented retry counters. [e2e-llm-inference-service] [e2e-llm-inference-service] :param response: A response object, or None, if the server did not [e2e-llm-inference-service] return a response. [e2e-llm-inference-service] :type response: :class:`~urllib3.response.BaseHTTPResponse` [e2e-llm-inference-service] :param Exception error: An error encountered during the request, or [e2e-llm-inference-service] None if the response was received successfully. [e2e-llm-inference-service] [e2e-llm-inference-service] :return: A new ``Retry`` object. [e2e-llm-inference-service] """ [e2e-llm-inference-service] if self.total is False and error: [e2e-llm-inference-service] # Disabled, indicate to re-raise the error. [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] [e2e-llm-inference-service] total = self.total [e2e-llm-inference-service] if total is not None: [e2e-llm-inference-service] total -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] connect = self.connect [e2e-llm-inference-service] read = self.read [e2e-llm-inference-service] redirect = self.redirect [e2e-llm-inference-service] status_count = self.status [e2e-llm-inference-service] other = self.other [e2e-llm-inference-service] cause = "unknown" [e2e-llm-inference-service] status = None [e2e-llm-inference-service] redirect_location = None [e2e-llm-inference-service] [e2e-llm-inference-service] if error and self._is_connection_error(error): [e2e-llm-inference-service] # Connect retry? [e2e-llm-inference-service] if connect is False: [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif connect is not None: [e2e-llm-inference-service] connect -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error and self._is_read_error(error): [e2e-llm-inference-service] # Read retry? [e2e-llm-inference-service] if read is False or method is None or not self._is_method_retryable(method): [e2e-llm-inference-service] raise reraise(type(error), error, _stacktrace) [e2e-llm-inference-service] elif read is not None: [e2e-llm-inference-service] read -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif error: [e2e-llm-inference-service] # Other retry? [e2e-llm-inference-service] if other is not None: [e2e-llm-inference-service] other -= 1 [e2e-llm-inference-service] [e2e-llm-inference-service] elif response and response.get_redirect_location(): [e2e-llm-inference-service] # Redirect retry? [e2e-llm-inference-service] if redirect is not None: [e2e-llm-inference-service] redirect -= 1 [e2e-llm-inference-service] cause = "too many redirects" [e2e-llm-inference-service] response_redirect_location = response.get_redirect_location() [e2e-llm-inference-service] if response_redirect_location: [e2e-llm-inference-service] redirect_location = response_redirect_location [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] else: [e2e-llm-inference-service] # Incrementing because of a server error like a 500 in [e2e-llm-inference-service] # status_forcelist and the given method is in the allowed_methods [e2e-llm-inference-service] cause = ResponseError.GENERIC_ERROR [e2e-llm-inference-service] if response and response.status: [e2e-llm-inference-service] if status_count is not None: [e2e-llm-inference-service] status_count -= 1 [e2e-llm-inference-service] cause = ResponseError.SPECIFIC_ERROR.format(status_code=response.status) [e2e-llm-inference-service] status = response.status [e2e-llm-inference-service] [e2e-llm-inference-service] history = self.history + ( [e2e-llm-inference-service] RequestHistory(method, url, error, status, redirect_location), [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] new_retry = self.new( [e2e-llm-inference-service] total=total, [e2e-llm-inference-service] connect=connect, [e2e-llm-inference-service] read=read, [e2e-llm-inference-service] redirect=redirect, [e2e-llm-inference-service] status=status_count, [e2e-llm-inference-service] other=other, [e2e-llm-inference-service] history=history, [e2e-llm-inference-service] ) [e2e-llm-inference-service] [e2e-llm-inference-service] if new_retry.is_exhausted(): [e2e-llm-inference-service] reason = error or ResponseError(cause) [e2e-llm-inference-service] > raise MaxRetryError(_pool, url, reason) from reason # type: ignore[arg-type] [e2e-llm-inference-service] E urllib3.exceptions.MaxRetryError: HTTPSConnectionPool(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Max retries exceeded with url: /api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload (Caused by NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Failed to resolve 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)")) [e2e-llm-inference-service] [e2e-llm-inference-service] ../../python/kserve/.venv/lib64/python3.11/site-packages/urllib3/util/retry.py:519: MaxRetryError [e2e-llm-inference-service] ------------------------------ Captured log setup ------------------------------ [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig router-managed-prestop-hook-tes-667dca99 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig router-managed-prestop-hook-tes-667dca99 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig router-managed-prestop-hook-tes-667dca99 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig workload-single-cpu-prestop-hoo-66c41342 in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig workload-single-cpu-prestop-hoo-66c41342 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig workload-single-cpu-prestop-hoo-66c41342 [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1586 Checking LLMInferenceServiceConfig model-fb-opt-125m-prestop-hook-584a80fb in namespace kserve-ci-e2e-test [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1612 Resource not found, creating LLMInferenceServiceConfig model-fb-opt-125m-prestop-hook-584a80fb [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1622 ✓ Successfully created LLMInferenceServiceConfig model-fb-opt-125m-prestop-hook-584a80fb [e2e-llm-inference-service] ------------------------------ Captured log call ------------------------------- [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [test_prestop_hook] [2026-07-02T14:28:54.980321] start - args=(), kwargs={'test_case': TestCase(base_refs=['router-managed', 'workload-single-cpu', 'model-fb-opt-125m'], prompt='KServe is a', service_name='prestop-hook-test', endpoint='/v1/completions', max_tokens=20, payload_formatter=None, response_assertion=, wait_timeout=900, response_timeout=60, extra_headers=None, url_getter=None, expected_gateway=None, before_test=[], after_test=[], peers=[], llm_service={'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': None, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'prestop-hook-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-prestop-hook-tes-667dca99'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-prestop-hoo-66c41342'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-prestop-hook-584a80fb'}]}, [e2e-llm-inference-service] 'status': None}, model_name='facebook/opt-125m')} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:fixtures.py:1637 No HTTP proxy configured for k8s client [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [create_llmisvc] [2026-07-02T14:28:54.993065] start - args=(, {'api_version': 'serving.kserve.io/v1alpha1', [e2e-llm-inference-service] 'kind': 'LLMInferenceService', [e2e-llm-inference-service] 'metadata': {'annotations': None, [e2e-llm-inference-service] 'creation_timestamp': None, [e2e-llm-inference-service] 'deletion_grace_period_seconds': None, [e2e-llm-inference-service] 'deletion_timestamp': None, [e2e-llm-inference-service] 'finalizers': None, [e2e-llm-inference-service] 'generate_name': None, [e2e-llm-inference-service] 'generation': None, [e2e-llm-inference-service] 'labels': None, [e2e-llm-inference-service] 'managed_fields': None, [e2e-llm-inference-service] 'name': 'prestop-hook-test', [e2e-llm-inference-service] 'namespace': 'kserve-ci-e2e-test', [e2e-llm-inference-service] 'owner_references': None, [e2e-llm-inference-service] 'resource_version': None, [e2e-llm-inference-service] 'self_link': None, [e2e-llm-inference-service] 'uid': None}, [e2e-llm-inference-service] 'spec': {'baseRefs': [{'name': 'router-managed-prestop-hook-tes-667dca99'}, [e2e-llm-inference-service] {'name': 'workload-single-cpu-prestop-hoo-66c41342'}, [e2e-llm-inference-service] {'name': 'model-fb-opt-125m-prestop-hook-584a80fb'}]}, [e2e-llm-inference-service] 'status': None}), kwargs={} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:43 [create_llmisvc] [2026-07-02T14:28:55.079117] end - ✅ in 0.086s [e2e-llm-inference-service] INFO e2e.llmisvc.logging:logging.py:34 [wait_for_workload_pod_running] [2026-07-02T14:28:55.079274] start - args=('kserve-ci-e2e-test', 'prestop-hook-test'), kwargs={'timeout_seconds': 900} [e2e-llm-inference-service] INFO e2e.llmisvc.logging:test_llm_inference_service.py:1222 Waiting: No Running pods found for app.kubernetes.io/name=prestop-hook-test,kserve.io/component=workload in kserve-ci-e2e-test [e2e-llm-inference-service] assert [] [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=2, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ConnectionResetError(104, 'Connection reset by peer')': /api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=1, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ConnectTimeoutError(, 'Connection to a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com timed out. (connect timeout=None)')': /api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload [e2e-llm-inference-service] WARNING urllib3.connectionpool:connectionpool.py:868 Retrying (Retry(total=0, connect=None, read=None, redirect=None, status=None)) after connection broken by 'NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Failed to resolve 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)")': /api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [wait_for_workload_pod_running] [2026-07-02T14:42:04.879981] end - ❌ 789.801s: HTTPSConnectionPool(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Max retries exceeded with url: /api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload (Caused by NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Failed to resolve 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)")) [e2e-llm-inference-service] ERROR e2e.llmisvc.logging:logging.py:48 [test_prestop_hook] [2026-07-02T14:42:04.880212] end - ❌ 789.900s: HTTPSConnectionPool(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Max retries exceeded with url: /api/v1/namespaces/kserve-ci-e2e-test/pods?labelSelector=app.kubernetes.io%2Fname%3Dprestop-hook-test%2Ckserve.io%2Fcomponent%3Dworkload (Caused by NameResolutionError("HTTPSConnection(host='a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com', port=6443): Failed to resolve 'a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com' ([Errno -2] Name or service not known)")) [e2e-llm-inference-service] =============================== warnings summary =============================== [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-precise-prefix-cache-inline-config-workload-llmd-simulator-kvcache] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-with-gateway-ref-router-with-managed-route-model-fb-opt-125m-workload-llmd-simulator] [e2e-llm-inference-service] /workspace/source/python/kserve/.venv/lib64/python3.11/site-packages/pytest_asyncio/plugin.py:761: DeprecationWarning: The event_loop fixture provided by pytest-asyncio has been redefined in [e2e-llm-inference-service] /workspace/source/test/e2e/conftest.py:43 [e2e-llm-inference-service] Replacing the event_loop fixture with a custom implementation is deprecated [e2e-llm-inference-service] and will lead to errors in the future. [e2e-llm-inference-service] If you want to request an asyncio event loop with a scope other than function [e2e-llm-inference-service] scope, use the "scope" argument to the asyncio mark when marking the tests. [e2e-llm-inference-service] If you want to return different types of event loops, use the event_loop_policy [e2e-llm-inference-service] fixture. [e2e-llm-inference-service] [e2e-llm-inference-service] warnings.warn( [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-precise-prefix-cache-inline-config-workload-llmd-simulator-kvcache] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator0] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator1] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator2] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf0] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-with-gateway-ref-router-with-managed-route-model-fb-opt-125m-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m-with-lora-hf1] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-custom-route-timeout-scheduler-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-with-refs-scheduler-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-pd-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-pd-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-custom-route-timeout-pd-scheduler-managed-workload-pd-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-pvc] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-with-refs-pd-scheduler-managed-workload-pd-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service_stop.py::test_llm_stop_feature[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service_stop.py:40: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-no-scheduler-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-inline-config-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-model-qwen2.5-0.5b] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_tls.py::test_llm_tls_resources[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] llmisvc/test_llm_tls.py:93: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-configmap-ref-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-replicas-workload-llmd-simulator] [e2e-llm-inference-service] llmisvc/test_llm_inference_service.py:245: PytestWarning: The test is marked with '@pytest.mark.asyncio' but it is not an async function. Please remove the asyncio mark. If the test is not marked explicitly, check for global marks applied via 'pytestmark'. [e2e-llm-inference-service] @pytest.mark.llminferenceservice [e2e-llm-inference-service] [e2e-llm-inference-service] -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html [e2e-llm-inference-service] ---------- generated xml file: /workspace/artifacts-dir/junit_e2e.xml ---------- [e2e-llm-inference-service] --------------------------------- JSON report ---------------------------------- [e2e-llm-inference-service] report saved to: /workspace/artifacts-dir/e2e_results.json [e2e-llm-inference-service] =========================== short test summary info ============================ [e2e-llm-inference-service] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_multi_node-router-managed-workload-simulated-dp-ep-cpu-model-fb-opt-125m] [e2e-llm-inference-service] FAILED llmisvc/test_llm_lora_adapters.py::test_llm_with_lora_adapters[cluster_cpu-single-lora-adapter-hf] [e2e-llm-inference-service] FAILED llmisvc/test_llm_lora_adapters.py::test_llm_with_lora_adapters[cluster_cpu-multiple-lora-adapters] [e2e-llm-inference-service] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-model-qwen2.5-0.5b] [e2e-llm-inference-service] FAILED llmisvc/test_llm_tls.py::test_llm_tls_resources[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-configmap-ref-workload-llmd-simulator] [e2e-llm-inference-service] FAILED llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-replicas-workload-llmd-simulator] [e2e-llm-inference-service] FAILED llmisvc/test_prestop_hook.py::test_prestop_hook[cluster_cpu-cluster_single_node-router-managed-workload-single-cpu-model-fb-opt-125m] [e2e-llm-inference-service] ERROR llmisvc/test_llm_inference_service.py::test_llm_inference_service[cluster_cpu-cluster_single_node-router-managed-scheduler-with-custom-template-workload-llmd-simulator] [e2e-llm-inference-service] ERROR llmisvc/test_rolling_upgrade.py::test_rolling_upgrade_coordination[cluster_cpu-cluster_single_node-router-managed-workload-llmd-simulator-model-fb-opt-125m] [e2e-llm-inference-service] !!!!!!!!!!!!!!!!!!!!!!!!!! stopping after 10 failures !!!!!!!!!!!!!!!!!!!!!!!!!! [e2e-llm-inference-service] !!!!!!!!!!!! xdist.dsession.Interrupted: stopping after 5 failures !!!!!!!!!!!!! [e2e-llm-inference-service] = 8 failed, 29 passed, 3 skipped, 26 warnings, 2 errors in 6550.11s (1:49:10) == [must-gather] [must-gather ] OUT 2026-07-02T14:42:07.185654288Z Using must-gather plug-in image: quay.io/modh/must-gather:rhoai-2.24 [must-gather] When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information: [must-gather] error getting cluster version: Get "https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/config.openshift.io/v1/clusterversions/version": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host [must-gather] ClusterID: [must-gather] ClientVersion: 4.21.10 [must-gather] ClusterVersion: Installing "" for : [must-gather] error getting cluster operators: Get "https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/config.openshift.io/v1/clusteroperators": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host [must-gather] ClusterOperators: [must-gather] clusteroperators are missing [must-gather] [must-gather] [must-gather] [must-gather] [must-gather] Error running must-gather collection: [must-gather] creating temp namespace: Post "https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api/v1/namespaces": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host [must-gather] [must-gather] Falling back to `oc adm inspect clusterversion.v1.config.openshift.io,clusteroperators.v1.config.openshift.io` to collect basic cluster types. [must-gather] E0702 14:42:07.223953 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] E0702 14:42:07.231975 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] E0702 14:42:07.240173 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] E0702 14:42:07.248573 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] E0702 14:42:07.259091 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] E0702 14:42:07.264722 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] E0702 14:42:07.269696 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] E0702 14:42:07.278961 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] error completing cluster type inspection: error running backup collection: Get "https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host [must-gather] Falling back to `oc adm inspect namespace/openshift-cluster-version` to collect basic cluster named resources. [must-gather] E0702 14:42:07.288609 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] E0702 14:42:07.296782 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] E0702 14:42:07.307321 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] E0702 14:42:07.312655 24 memcache.go:265] "Unhandled Error" err="couldn't get current server API group list: Get \"https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s\": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host" [must-gather] error completing cluster named resource inspection: error running backup collection: Get "https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api?timeout=32s": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host [must-gather] [must-gather] [must-gather] Reprinting Cluster State: [must-gather] When opening a support case, bugzilla, or issue please include the following summary data along with any other requested information: [must-gather] error getting cluster version: Get "https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/config.openshift.io/v1/clusterversions/version": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host [must-gather] ClusterID: [must-gather] ClientVersion: 4.21.10 [must-gather] ClusterVersion: Installing "" for : [must-gather] error getting cluster operators: Get "https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/apis/config.openshift.io/v1/clusteroperators": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host [must-gather] ClusterOperators: [must-gather] clusteroperators are missing [must-gather] [must-gather] [must-gather] error: creating temp namespace: Post "https://a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com:6443/api/v1/namespaces": dial tcp: lookup a67b4de1572174690a9723adeab422e0-4b74d5631435c099.elb.us-east-1.amazonaws.com on 172.30.0.10:53: no such host [git-push-artifacts] WORK_DIR: /workspace/odh-ci-artifacts [git-push-artifacts] REPO_PATH: opendatahub-io/odh-build-metadata [git-push-artifacts] REPO_BRANCH: ci-artifacts [git-push-artifacts] SPARSE_FILE_PATH: test-artifacts/docs [git-push-artifacts] SOURCE_PATH: /workspace/artifacts-dir [git-push-artifacts] DEST_PATH: test-artifacts/kserve-group-test-ffjdc [git-push-artifacts] ALWAYS_PASS: false [git-push-artifacts] configuring gh token [git-push-artifacts] taking github token from Konflux bot [git-push-artifacts] Initialized empty Git repository in /workspace/odh-ci-artifacts/.git/ [git-push-artifacts] Using partial fetch with sparse checkout for: test-artifacts/docs [git-push-artifacts] From https://github.com/opendatahub-io/odh-build-metadata [git-push-artifacts] * branch ci-artifacts -> FETCH_HEAD [git-push-artifacts] * [new branch] ci-artifacts -> origin/ci-artifacts [git-push-artifacts] Already on 'ci-artifacts' [git-push-artifacts] branch 'ci-artifacts' set up to track 'origin/ci-artifacts'. [git-push-artifacts] TASK_NAME=kserve-group-test-ffjdc-e2e-llm-inference-service [git-push-artifacts] PIPELINERUN_NAME=kserve-group-test-ffjdc [git-push-artifacts] From https://github.com/opendatahub-io/odh-build-metadata [git-push-artifacts] * branch ci-artifacts -> FETCH_HEAD [git-push-artifacts] Already up to date. [git-push-artifacts] -rw-r--r--. 1 root 1001540000 302188 Jul 2 14:42 /workspace/odh-ci-artifacts/test-artifacts/kserve-group-test-ffjdc/e2e-llm-inference-service.tar.gz [git-push-artifacts] [ci-artifacts 8b5b28f] Updating CI Artifacts in e2e-llm-inference-service [git-push-artifacts] 1 file changed, 0 insertions(+), 0 deletions(-) [git-push-artifacts] create mode 100644 test-artifacts/kserve-group-test-ffjdc/e2e-llm-inference-service.tar.gz [git-push-artifacts] From https://github.com/opendatahub-io/odh-build-metadata [git-push-artifacts] * branch ci-artifacts -> FETCH_HEAD [git-push-artifacts] Already up to date. [git-push-artifacts] To https://github.com/opendatahub-io/odh-build-metadata.git [git-push-artifacts] 0093d2d..8b5b28f ci-artifacts -> ci-artifacts [fail-if-needed] Failing pipeline because deploy-and-e2e step failed container step-fail-if-needed has failed : [{"key":"StartedAt","value":"2026-07-02T14:42:13.077Z","type":3}]